Audio Focus Selection Using Dominant Direction Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current parametric spatial audio systems require multiple pre-focused audio signals for flexible audio focus, leading to increased bit requirements for storage and transmission, which is inefficient.
Innovation Solution
The method involves capturing audio using at least three microphones, determining a dominant sound source direction, and selecting two microphones closest to this direction to create focused audio signals, reducing the number of bits needed for transmission by limiting focus directions based on the dominant sound source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple pre-focused audio signals are transmitted to enable flexible audio focus, then adaptability is improved, but the quantity of data to be stored and transmitted increases
Solution Approach 1:
The system performs preliminary beamforming to create a small set of pre-focused audio signals (typically two signals focused in opposite directions) before transmission. This preliminary action captures the essential spatial information needed for audio focus while avoiding the need to transmit all possible focused signals, thus resolving the contradiction between adaptability and data quantity.
Solution Approach 2:
The system changes the parameter of focus direction based on metadata indicating the dominant sound source direction. Instead of transmitting multiple fixed focused signals, the system transmits a small set of signals with parameters that can be adjusted at playback based on the metadata, reducing the quantity of data while maintaining adaptability.
2Adaptability or versatility
If beamforming and spatial filtering are used to achieve audio focus, then audio focus capability is improved, but the complexity of the system increases due to requiring knowledge of sound directions
Solution Approach 1:
The system introduces metadata as an intermediary that carries information about the dominant sound source direction. This metadata acts as a mediator between the captured audio signals and the beamforming process, providing the necessary directional information without requiring complex real-time analysis at the playback device, thus reducing system complexity while maintaining audio focus capability.
Data Source
AI summary
A method including: obtaining at least three microphone audio signals, wherein the microphone audio signals are associated with microphones with a location relative to an apparatus on which the microphones are located; analysing the at least three microphone audio signals to determine at least one metadata directional parameter; generating a first audio signal and a second audio signal based on at least one of the at least three microphone audio signals and the at least one metadata directional parameter; and outputting and/or storing the first audio signal, the second audio signal and the at least one metadata directional parameter, such that the first audio signal, the second audio signal, and the at least one metadata directional parameter enable a generation of an output audio signal with an adjustable audio focusing.


