Spatial Audio Rendering via Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio technologies face challenges in accurately rendering audio signals based on user device position within a virtual space, as they struggle to fully separate individual sound sources from composite audio signals, leading to degraded audio quality and immersion issues when movement is not seamlessly translated into audio scene changes.

Innovation Solution

The system employs spatial audio capture apparatuses to receive composite audio signals, identifies user device positions, and renders audio differently based on successful separation of individual sound sources within predetermined areas, using measures like correlation between composite and reference signals to determine separation success, enabling volumetric audio rendering and six degrees-of-freedom movement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If individual audio signals are separated from composite audio signals to enable accurate spatial rendering, then audio quality and immersion are improved, but separation success rate deteriorates when multiple sound sources are present

Engineering Contradiction:
Improveaudio rendering accuracyVSAvoidseparation success rate
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system attempts to separate individual audio signals from composite signals when possible (partial action), but does not require complete separation success for all sound sources. When separation succeeds for some sources, those are rendered with high accuracy while others fall back to alternative methods, achieving useful audio rendering without requiring perfect separation in all cases.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically changes rendering parameters based on separation success. When separation succeeds, parameters are set for high-accuracy individual source rendering; when separation fails, parameters switch to alternative rendering methods. This parameter adaptation resolves the contradiction by adjusting the approach based on actual signal conditions.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If volumetric audio rendering is implemented with six degrees-of-freedom movement tracking, then user immersion is improved, but device complexity increases

Engineering Contradiction:
Improvespatial rendering capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the virtual space into multiple predetermined areas, each with its own spatial audio capture apparatus. This segmentation allows complex volumetric rendering to be implemented in a modular way, where each area can be processed independently, reducing overall system complexity while maintaining immersive capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements a universal rendering framework that handles multiple types of audio signals (separated individual sources and composite signals) and multiple user movements (six degrees-of-freedom) through a single integrated apparatus. This multi-functionality achieves high adaptability without proportionally increasing complexity, as the same infrastructure handles diverse rendering scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If audio rendering is adjusted based on separation success to maintain quality, then audio consistency is improved, but processing time increases

Engineering Contradiction:
Improveaudio quality consistencyVSAvoidrendering processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs audio signal separation and success evaluation in advance before final rendering. By preprocessing the composite signals and determining separation outcomes beforehand, the system avoids time-consuming processing during actual rendering, thus maintaining audio quality consistency while reducing real-time processing delays.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11631422B2Methods, apparatuses and computer programs relating to spatial audio
Publication Date: 2023.04.18 NOKIA TECHNOLOGIES OY
  • US11631422B2 patent drawing
  • US11631422B2 patent drawing
  • US11631422B2 patent drawing

AI summary

An apparatus is disclosed, configured to receive, from first and second spatial audio capture apparatuses, respective first and second composite audio signals comprising components derived from one or more sound sources in a capture space. The apparatus is further configured to identify a position of a user device corresponding to one of first and second areas respectively associated with the positions of the first and second spatial audio capture apparatuses, and to render audio representing the one or more sound sources to the user device, the rendering being performed differently dependent on, for the spatial audio capture apparatus associated with the identified first or second area, whether or not individual audio signals from each of the one or more sound sources can be successfully separated from its composite signal.