Audio Signal Type Classification for Spatial Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio decoding technologies face challenges in rendering spatial audio signals effectively across different transport audio signal types, leading to audio quality deterioration when applying processing techniques designed for specific signal types to others.

Innovation Solution

The proposed solution involves determining the type of transport audio signals based on parameters such as energy ratios and spatial metadata, and then processing the signals accordingly to convert them into appropriate formats like Ambisonics or multichannel audio, using techniques like linear and parametric rendering to minimize artefacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If processing techniques designed for specific signal types are applied to other signal types, then processing efficiency is improved, but audio quality deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system changes the parameter of signal type identification to select appropriate processing techniques. By detecting whether the transport audio signal is spaced audio or downmix audio, the system dynamically adjusts the processing parameters (linear rendering vs parametric rendering) to maintain optimal audio quality for each signal type while preserving processing efficiency.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The processing system is segmented into different processing paths based on signal type. Spaced audio signals follow one processing path (linear rendering), while downmix audio signals follow another path (parametric rendering). This segmentation ensures each signal type receives specialized processing, preventing quality deterioration while maintaining overall system efficiency.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If a single processing method is used for all transport audio signal types, then device complexity is reduced, but audio quality deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidaudio quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The processing system dynamically adapts its complexity based on the input signal type. A type detection mechanism determines whether the signal is spaced audio or downmix audio, and the system dynamically selects the appropriate processing method. This dynamic approach allows the system to use simple processing when possible while applying more complex specialized processing only when needed, balancing device complexity with audio quality.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If processing is tailored to specific signal types, then audio quality is improved, but device complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by detecting the signal type before applying processing. The type detection step identifies whether the transport audio signal is spaced audio or downmix audio, allowing the system to pre-select the appropriate processing path. This preliminary classification prevents unnecessary complex processing from being applied to signals that don't require it, thereby improving audio quality while minimizing the actual processing complexity executed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240357304A1Sound Field Related Rendering
Publication Date: 2024.10.24 NOKIA TECHNOLOGIES OY
  • US20240357304A1 patent drawing
  • US20240357304A1 patent drawing
  • US20240357304A1 patent drawing

AI summary

An apparatus including circuitry configured to: obtain at least two audio signals; determine a type of the at least two audio signals; process the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals.