Spatial Audio Rendering with Notional Points-of-View

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio rendering technologies limit user immersion in spatial audio experiences due to restricted degrees of freedom, particularly when physical movement and orientation changes are constrained, leading to reduced engagement and realism in audio environments.

Innovation Solution

An apparatus and method that switch between two modes: a first mode where virtual sound scenes are rendered based on the user's current point-of-view, and a second mode where a sequence of notional points-of-view is determined to render virtual sound scenes, allowing for varying trajectories and stylistic continuity across different audio content items, even with limited user movement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the system uses user's current point-of-view to render virtual sound scenes, then the audio rendering responds to user movement, but the user immersion is limited when physical movement is restricted

Engineering Contradiction:
Improveuser movement freedomVSAvoidaudio rendering quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent creates notional points-of-view that copy the essential characteristics of physical user movement through virtual trajectories. When physical movement is restricted, the system generates synthetic point-of-view sequences that replicate the intended audio experience, allowing users to access immersive spatial audio content even without physical displacement.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameter of point-of-view determination from purely user-based to a hybrid approach combining user input with algorithmically generated notional points-of-view. This parameter change enables the system to maintain audio rendering quality by supplementing limited user movement with computationally generated trajectory data.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the system generates notional points-of-view automatically, then user immersion is enhanced through varied trajectories, but the system complexity increases

Engineering Contradiction:
Improveaudio rendering qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent pre-generates notional points-of-view and trajectories before actual audio rendering occurs. By calculating and storing these virtual movement paths in advance, the system reduces real-time computational complexity while maintaining high audio rendering quality. The preliminary action of creating notional trajectories allows the system to quickly switch between different audio perspectives without complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system uses limited user movement, then the device requirements are reduced, but the user engagement and realism are reduced

Engineering Contradiction:
Improvedevice adaptabilityVSAvoidaudio experience quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces notional points-of-view as an intermediary between limited user movement and the desired immersive audio experience. These notional points act as virtual proxies that translate minimal user input into rich spatial audio trajectories, allowing low-capability devices to deliver high-quality audio experiences by mediating between user constraints and audio rendering requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11089426B2Apparatus, method or computer program for rendering sound scenes defined by spatial audio content to a user
Publication Date: 2021.08.10 NOKIA TECHNOLOGIES OY
  • US11089426B2 patent drawing
  • US11089426B2 patent drawing
  • US11089426B2 patent drawing

AI summary

An apparatus comprising means for:in a first mode rendering sound scenes defined by a spatial audio content to a user, wherein a current sound scene is selected by a current point-of-view of the user; andin a second mode,automatically determining, at least in part, a sequence of notional points-of-view of the user in dependence upon the spatial audio content; andrendering sound scenes defined by the spatial audio content to a user, wherein a sequence of sound scenes are selected by the sequence of notional points-of-view of the user.