Spatial Audio Rendering via Metadata-Driven Ambisonics Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spatial audio rendering technologies face challenges in efficiently simulating real-world acoustic environments in virtual settings, particularly in terms of rendering multiple sound sources and dynamic scenes, due to high computational requirements and limitations in rendering quality on devices with limited computing power.

Innovation Solution

A system and method for spatial audio rendering that uses metadata to determine rendering parameters, including acoustic environment, listener, and sound source information, to process audio signals using Ambisonics and spatial impulse responses, enabling efficient encoding and decoding of audio signals to simulate realistic sound propagation and environment acoustics, even on devices with weak computing power.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If ray tracing method is used to simulate sound propagation, then rendering quality is improved, but computational requirements increase significantly

Engineering Contradiction:
Improverendering qualityVSAvoidcomputational requirements
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The patent segments the sound propagation simulation into multiple components: direct sound, early reflections, and late reverberation. Each component is processed separately using different computational methods, allowing high-quality rendering where needed while reducing overall computational load compared to full ray tracing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the acoustic environment into a parameterized representation using metadata (room dimensions, material properties, sound source positions) and processes audio signals through parameter-based spatial encoding (Ambisonics order, rendering parameters) rather than computationally intensive geometric ray tracing.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If full spatial audio processing is performed, then spatial accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespatial accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing by encoding audio signals into spatial representations (Ambisonics) and pre-calculating rendering parameters based on metadata before real-time playback. This allows fast decoding and rendering while maintaining spatial accuracy, as the computationally intensive work is done in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic spatial audio rendering where the level of processing adapts to the scenario. Full spatial processing is applied only when needed (e.g., multiple sound sources, complex environments), while simpler scenarios use reduced processing, optimizing the balance between spatial accuracy and processing time.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240244388A1System and method for spatial audio rendering, and electronic device
Publication Date: 2024.07.18 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20240244388A1 patent drawing
  • US20240244388A1 patent drawing
  • US20240244388A1 patent drawing

AI summary

The present disclosure relates to a method for spatial audio rendering. The method comprises: on the basis of metadata, determining a parameter for spatial audio rendering, wherein the metadata comprises at least some information among acoustic environment information, listener spatial information and sound source spatial information, and the parameter for spatial audio rendering indicates a characteristic of sound propagation in a scene in which a listener is located; on the basis of the parameter for spatial audio rendering, processing an audio signal of a sound source, so as to obtain an encoded audio signal; and performing spatial decoding on the encoded audio signal, so as to obtain a decoded audio signal.