Spatial Audio Previewing via Sound Source Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In spatial audio, users face challenges in navigating and finding specific scenes due to the passive nature of standard audio tracks, where they can only search through time, whereas spatial audio requires searching through multiple dimensions of space and time, making it difficult to locate desired content.

Innovation Solution

A method for previewing spatial audio scenes using user-selected sound sources and related contextual sound sources, allowing for efficient browsing and selection by rendering an audio preview that includes the selected sound source and contextual sound sources, enabling users to make informed decisions and filter content effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If spatial audio scenes with multiple sound sources are rendered, then the audio content provides rich spatial information and immersive experience, but it becomes difficult for users to navigate and find specific scenes

Engineering Contradiction:
Improvespatial audio experienceVSAvoidcontent navigation
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent segments the complex spatial audio scene into individual sound source components. Each sound source can be independently selected, previewed, and manipulated. This segmentation allows users to navigate through audio content by selecting specific sound sources rather than searching through the entire spatial audio scene, thereby improving ease of operation while maintaining the rich spatial audio experience.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If all sound sources in a spatial audio scene are included in the preview, then the preview provides complete information, but the preview becomes complex and overwhelming for users

Engineering Contradiction:
Improveinformation completenessVSAvoidpreview complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts and highlights only the user-selected sound source in the audio preview, separating it from other sound sources in the spatial audio scene. This extraction approach provides users with focused information about the selected sound source without the complexity of the complete scene, allowing informed decisions while maintaining simplicity in the preview interface.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If users can search through multiple dimensions of space and time in spatial audio, then they can locate desired content more precisely, but the search complexity increases significantly

Engineering Contradiction:
Improvescene location precisionVSAvoidsearch complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces sound source selection as an intermediary mechanism between the user and the multi-dimensional spatial audio search space. Instead of directly navigating through complex space and time dimensions, users select sound sources as intermediaries that represent specific locations and time points in the spatial audio scene. This intermediary approach maintains precise scene location capability while significantly reducing search complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3570566B1Previewing spatial audio scenes comprising multiple sound sources
Publication Date: 2022.12.28 NOKIA TECHNOLOGIES OY
  • EP3570566B1 patent drawingFigure 1~3
  • EP3570566B1 patent drawingFigure 4A~6F
  • EP3570566B1 patent drawingFigure 7~8B

AI summary

An apparatus comprising means for: in response to user input, selecting at least one sound source of a spatial audio scene, comprising multiple sound sources, the spatial audio scene being defined by spatial audio content; selecting at least one related contextual sound source based on the at least one selected sound source; and causing rendering of an audio preview, representing the spatial audio content, that can be selected by a user, wherein the audio preview comprises a mix of sound sources including at least the at least one selected sound source and the at least one related contextual sound source but not all of the multiple sound sources of the spatial audio scene, and wherein selection of the audio preview causes an operation on at least the selected sound source.