Spatial Audio Rendering via Semantic Signal Decomposition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio processing methods face challenges in achieving high perceptual quality for noise-like signals and ambience materials, such as applause and natural environments, due to unsatisfactory quality or high computational complexity, particularly in decorrelating and up-mixing processes.

Innovation Solution

The approach involves decomposing audio signals into foreground and background components, which are then processed separately based on their semantic properties, allowing for adaptive spatial rendering and decorrelation, thereby improving perceptual quality while maintaining moderate computational costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional decorrelation methods are used for noise-like signals and ambience materials, then spatial audio rendering can be achieved, but perceptual quality becomes unsatisfactory

Engineering Contradiction:
Improveperceptual qualityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The input audio signal is decomposed into multiple semantic signal components (e.g., foreground events like handclaps and background ambience). Each component is processed independently through separate rendering paths, allowing tailored spatial processing for different signal types while maintaining overall system efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different rendering characteristics are applied to different signal components based on their semantic properties. Foreground components receive one type of spatial processing while background components receive another, optimizing perceptual quality for each component's specific characteristics rather than applying a uniform approach.

Inventive Principle:
Principle #3Local quality

2Reliability

If object-orientated approaches are used to model auditory events, then spatial audio quality improves, but computational complexity increases due to the number of auditory events to be processed

Engineering Contradiction:
Improvespatial audio qualityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal is segmented into semantic components that group similar auditory events together. Instead of processing each individual auditory event separately, the system processes groups of events with similar characteristics through shared rendering paths, significantly reducing computational complexity while maintaining spatial audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter representation from individual event-based modeling to semantic component-based modeling. By transforming the problem from processing numerous discrete auditory events to processing a smaller number of semantic signal components with distinct characteristics, computational complexity is reduced while preserving spatial rendering quality.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If strong decorrelation is applied to restore ambience sensation, then spatial immersion improves, but transient event quality degrades due to temporal smearing effects

Engineering Contradiction:
Improvespatial immersionVSAvoidtransient event quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The system separates transient events (foreground) from ambience signals (background) into different signal components. This segmentation allows applying different rendering characteristics to each component: mild processing for transients to preserve their sharpness and strong decorrelation for background ambience to enhance spatial immersion, thereby resolving the contradiction between the two quality requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different levels of decorrelation strength are applied locally to different signal components based on their semantic properties. Foreground transient components receive processing that preserves temporal precision, while background ambience components receive stronger decorrelation to enhance spatial immersion. This local differentiation resolves the contradiction by optimizing each component for its specific quality requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP2418877B1An apparatus for determining a spatial output multi-channel audio signal
Publication Date: 2015.09.09 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2418877B1 patent drawingFigure 1A
  • EP2418877B1 patent drawingFigure 1B
  • EP2418877B1 patent drawingFigure 2~3

AI summary

An apparatus (100) for determining a spatial output multi-channel audio signal based on an input audio signal and an input parameter. The apparatus (100) comprises a decomposer (110) for decomposing the input audio signal based on the input parameter to obtain a first decomposed signal and a second decomposed signal different from each other. Furthermore, the apparatus (100) comprises a renderer (110) for rendering the first decomposed signal to obtain a first rendered signal having a first semantic property and for rendering the second decomposed signal to obtain a second rendered signal having a second semantic property being different from the first semantic property. The apparatus (100) comprises a processor (130) for processing the first rendered signal and the second rendered signal to obtain the spatial output multi-channel audio signal.