Spatial Audio Metadata Encoding Using Image-Based Source Location
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies do not effectively utilize spatial metadata parameters in metadata assisted spatial audio (MASA) formats, particularly in encoding multi-microphone audio, leading to suboptimal spatial audio reproduction.
Innovation Solution
An apparatus and method that obtain image-based sound source location data from image analysis and encode it within spatial audio metadata parameters, adjusting parameters like direction index, spread coherence, and surround coherence to enhance spatial audio distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spatial metadata parameters are used in MASA format for encoding multi-microphone audio, then spatial audio reproduction quality is improved, but the utilization effectiveness of these parameters remains suboptimal
Solution Approach 1:
The patent transforms image-based sound source location data into spatial audio metadata parameters (direction index, spread coherence, surround coherence) by mapping visual spatial information to audio parameter space. This parameter transformation enables effective utilization of previously underused spatial metadata fields, improving spatial audio reproduction accuracy without information loss.
2Manufacturing precision
If image-based sound source location data is encoded into spatial audio metadata parameters, then spatial distribution accuracy is improved, but the encoding complexity increases
Solution Approach 1:
The patent introduces image-based sound source location data as an intermediary that bridges visual and audio modalities. This intermediary data structure enables accurate spatial distribution encoding by serving as a common representation that can be systematically transformed into spatial audio metadata parameters, managing encoding complexity through structured data transformation.
Solution Approach 2:
The patent performs preliminary extraction and structuring of sound source location data from images before the actual encoding process. By pre-processing visual data to identify and organize spatial information about sound sources, the system simplifies the subsequent encoding step into spatial audio metadata, reducing overall encoding complexity while maintaining high spatial distribution accuracy.
Data Source
AI summary
An apparatus comprising means for:obtaining image-based sound source location data from image analysis of one or more captured images;encoding multi-microphone audio as metadata assisted spatial audio comprising spatial audio metadata parameters;encoding the image-based sound source location data within one or more spatial audio metadata parameters of the metadata assisted spatial audio, wherein the one or more spatial audio metadata parameters encoding the image-based sound source location data is or are one or more spatial audio metadata parameters defining a spatial distribution of audio energy.


