Audio Object Classification Using Location Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems face challenges in efficiently classifying audio objects based on location metadata, particularly in real-time applications with many audio objects, leading to high computational complexity and potential audible artefacts.
Innovation Solution
A method that estimates the presence of dialog in audio objects using location metadata, assigning a confidence value to an object type parameter based on position, speed, acceleration, and elevation, allowing for selective dialog enhancement and clustering to reduce computational load and improve bitrate efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio objects are classified using traditional audio content analysis methods, then classification accuracy is improved, but computational complexity increases significantly
Solution Approach 1:
The patent segments the classification task by using location metadata as a preliminary filter to identify candidate audio objects that may contain dialog, rather than analyzing all audio objects. This divides the computational workload into two stages: initial filtering based on spatial position and subsequent detailed analysis only for selected objects, thereby reducing overall computational complexity while maintaining classification accuracy.
Solution Approach 2:
The patent introduces location metadata as an intermediary element that mediates between the audio content and the classification process. By using spatial position information as an intermediate filtering criterion, the system avoids direct complex analysis of all audio content, reducing computational load while preserving classification precision through the intermediate spatial selection step.
2Reliability
If dialog enhancement is applied to all audio objects, then dialog quality is improved, but bitrate efficiency deteriorates
Solution Approach 1:
The patent applies dialog enhancement selectively rather than uniformly across all audio objects. By using location metadata to identify which audio objects are likely to contain dialog (e.g., objects at specific spatial positions or with certain movement characteristics), the system applies enhancement only to those specific objects, thereby improving dialog quality where needed while avoiding unnecessary processing and maintaining bitrate efficiency.
Solution Approach 2:
The patent changes the parameter used for selecting audio objects for enhancement from traditional content-based features to location metadata parameters. This parameter change enables more efficient selection of candidates for dialog enhancement, improving the trade-off between dialog quality and bitrate efficiency by using spatial information as the selection criterion instead of computationally intensive content analysis.
3Stability of the object's composition
If clustering is performed on all audio objects, then audio scene organization is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary filtering using location metadata before conducting clustering operations. By pre-selecting audio objects that are likely to contain dialog based on their spatial characteristics, the system reduces the number of objects that need to be clustered, thereby maintaining audio scene organization while significantly reducing processing time required for the clustering operation.
Data Source
AI summary
Methods (700, 800, 900), systems (200, 300, 400, 500, 600) and computer program products are provided. Location metadata (620) associated with an audio object is received (801). The location metadata defines a position of the audio object in an audio scene. It is estimated (630, 802), based on the location metadata, whether the audio object includes dialog. A value representative of a result of the estimation is assigned (803) to an object type parameter (231). In some example embodiments, audio objects are selected (661, 662, 804) based on values of their respective of object type parameters. In some example embodiments, at least one of the selected audio objects is submitted to dialog enhancement (690, 807).


