Binaural Audio Rendering Using Metadata for Computational Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio signal processing technologies face challenges in efficiently rendering 3D audio signals with a smaller number of channels, particularly in binaural rendering, which requires accurate simulation of sound images and motion application without excessive computational resources.

Innovation Solution

The proposed audio signal processing method and device utilize metadata to determine the rendering of audio signals across multiple tracks, allowing for binaural rendering with fewer channels by inserting metadata that indicates sound levels, binaural effect levels, and motion application information, enabling efficient processing and output of 3D audio signals using a smaller number of channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If binaural rendering is performed with accurate sound image simulation, then the three-dimensionality and spatial accuracy are improved, but the computational complexity and processing resources increase significantly

Engineering Contradiction:
Improvesound image simulation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple tracks, with metadata embedded in specific tracks to guide rendering. This segmentation allows the system to process only relevant audio components with binaural rendering while handling other components more efficiently, reducing overall computational complexity while maintaining sound image accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Metadata indicating sound levels, binaural effect levels, and motion application information is inserted into the audio file before rendering. This preliminary action provides pre-computed guidance information that enables the renderer to make decisions without performing complex real-time analysis, thereby reducing computational complexity during the rendering process while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If a smaller number of channels are used for audio output, then the device compatibility and processing efficiency are improved, but the 3D audio quality and spatial resolution deteriorate

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidspatial resolution
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system changes the parameter of channel count from multiple channels to a smaller number of channels (e.g., stereo) while using metadata parameters (sound level, binaural effect level, motion application) to compensate for the reduced spatial information. This allows efficient processing with fewer channels while maintaining spatial resolution through intelligent parameter utilization.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Metadata acts as an intermediary that carries spatial and rendering information between the multi-channel audio source and the fewer-channel output. The metadata includes sound level information, binaural effect level information, and motion application information, which guide the rendering process to preserve spatial resolution even when output channels are reduced.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If metadata is embedded in audio files to guide rendering, then the rendering flexibility and adaptability are improved, but the file complexity and processing overhead increase

Engineering Contradiction:
Improverendering flexibilityVSAvoidfile complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The metadata structure is designed to be universal and multi-functional, serving multiple purposes: guiding binaural rendering, indicating sound levels, specifying binaural effect levels, and providing motion application information. This multi-functionality reduces the need for multiple separate data structures, thereby reducing file complexity while maintaining rendering flexibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10659904B2Method and device for processing binaural audio signal
Publication Date: 2020.05.19 GAUDI AUDIO LAB
  • US10659904B2 patent drawing
  • US10659904B2 patent drawing
  • US10659904B2 patent drawing

AI summary

An audio signal processing device for processing an audio signal may include a receiving unit configured to receive an audio file including the audio signal, a processor configured to simultaneously render a first audio signal component included in a first track of the audio file and a second audio signal component included in a second track of the audio file, and an output unit configured to output the rendered first audio signal component and the rendered second audio signal component.