Spatial Audio Augmentation Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio technologies face challenges in seamlessly integrating augmentation audio signals with spatial audio signals, particularly in immersive audio environments like virtual reality, where user movement affects audio rendering.

Innovation Solution

The proposed solution involves an apparatus and method that obtain spatial audio signals and augmentation audio signals, render them consistently with user movement, and mix them to generate an output audio signal. This is achieved by decoding spatial and augmentation audio signals from separate bit streams, applying user-locked or world-locked mixing modes, and adjusting gains based on user position and rotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If spatial audio signals are rendered consistently with user movement in immersive environments, then audio fidelity and immersion quality are improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improveaudio fidelityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The audio processing system is segmented into separate spatial audio signal processing path and augmentation audio signal processing path. The spatial audio decoder processes spatial audio signals with full 6DoF movement compensation, while the augmentation audio decoder handles augmentation signals with simplified processing. This segmentation allows high-fidelity spatial audio rendering without requiring the entire system to handle maximum processing complexity simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts the rendering approach based on the type of audio signal being processed. Spatial audio signals receive dynamic rendering that fully compensates for 6DoF user movement, while augmentation audio signals use a different rendering approach. This dynamic differentiation maintains audio fidelity for spatial content while managing overall processing complexity.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If separate bit streams are decoded for spatial and augmentation audio signals, then audio quality is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The spatial audio signal and augmentation audio signal are decoded from separate bit streams in parallel preliminary processing stages. By preparing both signals beforehand through separate decoding paths, the system avoids sequential processing delays and can mix the rendered outputs more efficiently, reducing overall processing time while maintaining high audio quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

After separate decoding and rendering of spatial and augmentation audio signals, the processed signals are merged in the mixing stage. This combining of separately processed high-quality signals achieves both goals: maintains audio quality through dedicated processing paths while reducing processing time through parallel operation before the final mix.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If user-locked or world-locked mixing modes are applied, then adaptability to different usage scenarios is improved, but device complexity increases

Engineering Contradiction:
Improvemixing mode flexibilityVSAvoidcontrol system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Different mixing modes (user-locked and world-locked) are applied locally to different audio signal components based on their specific requirements. Spatial audio signals can use one mixing mode while augmentation audio signals use another, allowing tailored processing for each signal type without requiring the entire system to support all modes simultaneously, thus managing complexity while maintaining versatility.

Inventive Principle:
Principle #3Local quality

4Manufacturing precision

If gains are adjusted based on user position and rotation, then audio realism and immersion are improved, but computational load increases

Engineering Contradiction:
Improveaudio realismVSAvoidcomputational energy
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system applies gain adjustment based on user position and rotation selectively to spatial audio signals where it most impacts realism, rather than uniformly to all audio content. This partial application of the computationally intensive gain adjustment process maintains audio realism for spatial content while reducing overall computational energy consumption by not applying the same level of processing to all audio signals.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12267665B2Spatial audio augmentation
Publication Date: 2025.04.01 NOKIA TECHNOLOGIES OY
  • US12267665B2 patent drawing
  • US12267665B2 patent drawing
  • US12267665B2 patent drawing

AI summary

An apparatus configured to: obtain at least one spatial audio signal, the at least one spatial audio signal comprising at least one audio signal and at least one spatial parameter; determine at least one first audio signal based, at least partially, on the at least one spatial audio signal and the at least one of the user position or orientation; obtain at least one augmentation audio signal, wherein the at least one augmentation audio signal has a different audio format than an audio format of the at least one spatial audio signal, wherein the at least one augmentation audio signal provides a different type of media content; determine at least one second audio signal based, at least partially, on at least a part of the at least one augmentation audio signal; and mix the at least one first audio signal and the at least one second audio signal.