Spatial Audio Metadata Encoding Using Image-Based Source Location

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies do not effectively utilize spatial metadata parameters in metadata assisted spatial audio (MASA) formats, particularly in encoding multi-microphone audio, leading to suboptimal spatial audio reproduction.

Innovation Solution

An apparatus and method that obtain image-based sound source location data from image analysis and encode it within spatial audio metadata parameters, adjusting parameters like direction index, spread coherence, and surround coherence to enhance spatial audio distribution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatial metadata parameters are used in MASA format for encoding multi-microphone audio, then spatial audio reproduction quality is improved, but the utilization effectiveness of these parameters remains suboptimal

Engineering Contradiction:
Improvespatial audio reproduction accuracyVSAvoidunderutilization of spatial metadata
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transforms image-based sound source location data into spatial audio metadata parameters (direction index, spread coherence, surround coherence) by mapping visual spatial information to audio parameter space. This parameter transformation enables effective utilization of previously underused spatial metadata fields, improving spatial audio reproduction accuracy without information loss.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If image-based sound source location data is encoded into spatial audio metadata parameters, then spatial distribution accuracy is improved, but the encoding complexity increases

Engineering Contradiction:
Improvespatial distribution accuracyVSAvoidencoding process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces image-based sound source location data as an intermediary that bridges visual and audio modalities. This intermediary data structure enables accurate spatial distribution encoding by serving as a common representation that can be systematically transformed into spatial audio metadata parameters, managing encoding complexity through structured data transformation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary extraction and structuring of sound source location data from images before the actual encoding process. By pre-processing visual data to identify and organize spatial information about sound sources, the system simplifies the subsequent encoding step into spatial audio metadata, reducing overall encoding complexity while maintaining high spatial distribution accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260089455A1Apparatus, method, computer program for encoding multi-microphone audio as metadata assisted spatial audio
Publication Date: 2026.03.26 NOKIA TECHNOLOGIES OY
  • US20260089455A1 patent drawing
  • US20260089455A1 patent drawing
  • US20260089455A1 patent drawing

AI summary

An apparatus comprising means for:obtaining image-based sound source location data from image analysis of one or more captured images;encoding multi-microphone audio as metadata assisted spatial audio comprising spatial audio metadata parameters;encoding the image-based sound source location data within one or more spatial audio metadata parameters of the metadata assisted spatial audio, wherein the one or more spatial audio metadata parameters encoding the image-based sound source location data is or are one or more spatial audio metadata parameters defining a spatial distribution of audio energy.