Spatial Audio Object Separation for Efficient Hybrid Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio encoding technologies fail to efficiently utilize synergies in processing different types of audio signals, such as microphone-array captured signals, loudspeaker signals, and Ambisonic signals, leading to inefficient encoding and potential artifacts.

Innovation Solution

A method and apparatus for spatial audio encoding that separates audio objects from a plurality of audio objects, determining a loudest energy proportion factor, and using threshold comparisons to identify and encode audio objects with either a hard or fade transition, combining these with other input audio formats to exploit signal synergies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If audio objects are processed separately in a separate processing chain, then processing simplicity is maintained, but encoding efficiency and signal synergies are not efficiently utilized

Engineering Contradiction:
Improveencoding efficiencyVSAvoidprocessing chain complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the processing chains by integrating audio object separation and encoding within the main spatial audio encoding process. The encoder now processes both multi-channel audio signals and audio objects together, allowing synergies to be exploited while maintaining a unified processing architecture that does not significantly increase overall system complexity

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If all audio objects are encoded together with other input audio formats, then processing uniformity is maintained, but encoding efficiency for dominant audio objects is reduced

Engineering Contradiction:
Improveencoding efficiencyVSAvoidprocessing uniformity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments the audio objects into two groups: dominant audio objects that are separated and encoded individually to maximize efficiency, and non-dominant audio objects that are encoded together with other input audio formats. This segmentation is based on energy proportion factors and threshold comparisons, allowing the system to optimize encoding efficiency for important objects while maintaining processing uniformity for less critical ones

Inventive Principle:
Principle #1Segmentation

3Productivity

If audio objects are separated and encoded individually, then encoding efficiency for dominant objects is improved, but processing complexity and potential artifacts increase

Engineering Contradiction:
Improveencoding efficiencyVSAvoidartifacts
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent implements dynamic transition mechanisms between hard transitions and fade transitions when separating and encoding audio objects. The system adaptively selects the transition type based on the specific audio context and object characteristics, thereby improving encoding efficiency for dominant objects while dynamically controlling artifact generation to maintain audio quality

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250279103A1Separating spatial audio objects
Publication Date: 2025.09.04 NOKIA TECHNOLOGIES OY
  • US20250279103A1 patent drawing
  • US20250279103A1 patent drawing
  • US20250279103A1 patent drawing

AI summary

There is inter alia disclosed an apparatus for spatial audio encoding configured to: determine an audio object for separation (306) from a plurality of audio objects of an audio frame (1281); separate the audio object for separation (308) from the plurality of audio objects to provide a separated audio object (126) and at least one remaining audio object (124); encode the separated audio object with an audio object encoder; and encode the plurality of remaining audio objects together with another input audio format.