Spatial Audio Metadata Encoding Using Image-Based Source Location
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current implementations of metadata assisted spatial audio (MASA) formats, particularly in multi-microphone capture on UEs, underutilize coherence parameters such as spread and surround coherence, leading to inefficiencies in reproducing spatial audio accurately.
Innovation Solution
An apparatus and method that utilize image-based sound source location data to vary spatial audio metadata parameters, including direction indices, coherence parameters, and spread coherence parameters, to enhance the encoding of multi-microphone audio as metadata assisted spatial audio, thereby improving the accuracy and spatial distribution of audio energy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image-based sound source location data is integrated into spatial audio metadata parameters, then spatial distribution accuracy and coherence are improved, but device complexity and processing requirements increase
Solution Approach 1:
The system performs image analysis to obtain sound source location data before encoding the spatial audio metadata. By preparing the visual location data in advance, the encoding process can directly integrate this information into coherence parameters without adding complex real-time processing during audio encoding, thus improving spatial accuracy while managing device complexity.
Solution Approach 2:
The patent uses image-based sound source location data as an intermediary to inform and adjust spatial audio metadata parameters. This intermediary data bridge allows the system to translate visual spatial information into audio coherence parameters, improving spatial distribution accuracy without requiring direct complex interaction between audio and visual processing systems.
2Reliability
If coherence parameters are varied based on image-based sound source location data, then spatial audio reproduction quality is improved, but computational requirements and processing time increase
Solution Approach 1:
The system obtains sound source location data from image analysis before the audio encoding process. By preparing this spatial reference information in advance, the system can efficiently vary coherence parameters during encoding without requiring time-consuming real-time calculations, thus improving reproduction quality while minimizing processing time delays.
3Measurement precision
If spatial audio metadata parameters are enhanced with image-based location data, then directional sound coherence is improved, but data transmission bandwidth requirements increase
Solution Approach 1:
The patent applies image-based sound source location data selectively to specific spatial audio metadata parameters that benefit most from it, such as coherence parameters and direction indices. Rather than enhancing all metadata uniformly, this targeted approach improves directional sound coherence where needed while minimizing unnecessary data expansion, thus balancing precision improvement with bandwidth efficiency.
Data Source
Figure 1~4
Figure 5~8
Figure 9~10
AI summary
An apparatus comprising means for: obtaining image-based sound source location data from image analysis of one or more captured images; encoding multi-microphone audio as metadata assisted spatial audio comprising spatial audio metadata parameters; encoding the image-based sound source location data within one or more spatial audio metadata parameters of the metadata assisted spatial audio, wherein the one or more spatial audio metadata parameters encoding the image-based sound source location data is or are one or more spatial audio metadata parameters defining a spatial distribution of audio energy.