Spatial Audio Rendering with Parametric Interaction Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Designing virtual audio environments that react realistically to user interactions with multiple audio sources is inefficient due to the time and effort required to define and implement complex metadata-driven models.

Innovation Solution

A method and apparatus for rendering spatial audio signals in a selectable viewpoint audio environment, which involves receiving a selected listening position and orientation, detecting interactions with audio objects based on predefined criteria, and modifying audio objects to derive a spatial audio signal that reflects the interaction, allowing for more flexible and versatile reactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If complex metadata-driven models are used to define audio source interactions, then the virtual audio environment can react realistically to user interactions, but the time and effort required to design and implement these models increases significantly

Engineering Contradiction:
Improverealism of audio source interactionsVSAvoidtime and effort to design and implement interaction models
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses audio parametric representations that copy and generalize interaction patterns across multiple audio sources. Instead of defining separate complex metadata models for each audio source interaction, the system creates parametric templates that can be instantiated and modified for different sources, significantly reducing design time while maintaining realistic interaction behavior.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the approach from defining fixed metadata models to using adjustable audio parametric representations. These parameters can be dynamically modified to represent different interaction scenarios, allowing the same underlying model to adapt to various situations without requiring complete redesign, thus reducing both time and complexity.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If separate interaction models are defined for each audio source, then each source can react individually to user interactions, but the overall system complexity and implementation effort increases

Engineering Contradiction:
Improveindividual audio source reaction capabilityVSAvoidsystem complexity for managing multiple interaction models
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal audio parametric representation that can serve multiple audio sources simultaneously. This single versatile framework can model interactions for any number of audio sources with consistent behavior patterns, eliminating the need to create and manage separate complex models for each source while still allowing individual customization through parameter adjustment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments the interaction modeling into modular audio parametric representations that can be independently configured but collectively managed through a unified framework. This segmentation allows individual audio sources to have customized parameters while the overall system complexity is reduced through standardized management of these modular components.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11516615B2Audio processing
Publication Date: 2022.11.29 NOKIA TECHNOLOGIES OY
  • US11516615B2 patent drawing
  • US11516615B2 patent drawing
  • US11516615B2 patent drawing

AI summary

A method for rendering a spatial audio signal that represents a sound field in a selectable viewpoint audio environment that includes one or more audio objects associated with respective audio content and a respective position in the audio environment. The method includes receiving an indication of a selected listening position and orientation in the audio environment; detecting an interaction concerning a first audio object on basis of one or more predefined interaction criteria; modifying the first audio object and one or more further audio objects linked thereto; and deriving the spatial audio signal that includes at least audio content associated with the modified first audio object in a first spatial position of the sound field that corresponds to its position in the audio environment in relation to said selected listening position and orientation, and audio content associated with the modified one or more further audio objects.