Spatial Audio Directionality Adjustment via Unified Reference Frame

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for adjusting audio and video content after capture have limited options for modifying their spatial characteristics, such as perceived directionality, which restricts enhanced user perception.

Innovation Solution

A method and apparatus that receive spatial audio input, determine a direction of interest, and generate spatial audio and visual outputs based on this input, allowing for the adjustment of aural and visual directivity to align with the direction of interest, thereby enhancing user perception by synchronizing audio and visual cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If spatial audio content is captured with fixed spatial characteristics, then the audio recording is simple and straightforward, but the ability to adjust spatial characteristics after capture is limited

Engineering Contradiction:
Improveability to adjust spatial characteristicsVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by capturing audio signals in an uncompressed, high-resolution format with preserved spatial information during the recording phase. This allows flexible spatial manipulation to be performed later without degrading the original spatial characteristics, enabling both adaptability and maintaining quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent utilizes parameter changes by allowing dynamic adjustment of spatial audio parameters such as directionality, spatial position, and field of view after capture. This enables the same audio recording to be adapted for different applications and listening scenarios without re-recording.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If spatial audio processing is performed in real-time with high fidelity, then user perception is enhanced, but processing time and computational resources increase

Engineering Contradiction:
Improvespatial accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by processing only the essential spatial parameters that most significantly impact user perception. Rather than processing all possible audio parameters, the system focuses on key spatial characteristics such as direction and spatial position, achieving high perceptual fidelity with reduced computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If audio and video spatial characteristics are processed independently, then each modality can be optimized separately, but synchronization and coherence between audio and visual cues are difficult to achieve

Engineering Contradiction:
Improveprocessing flexibilityVSAvoidspatial coherence
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges audio and video processing by establishing a unified spatial reference frame that coordinates both modalities. This integration ensures that audio and visual cues are spatially coherent and synchronized, while still allowing independent optimization of each modality's processing parameters within the unified framework.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2837211B1Method, apparatus and computer program for generating an spatial audio output based on an spatial audio input
Publication Date: 2017.08.30 NOKIA TECHNOLOGIES OY
  • EP2837211B1 patent drawingFigure 1
  • EP2837211B1 patent drawingFigure 2
  • EP2837211B1 patent drawingFigure 3

AI summary

A method, apparatus and computer program for: receiving a spatial audio input; determining a direction of interest from the spatial audio input; and generating a spatial audio output dependent on the spatial audio input and the direction of interest.