Audio Object Classification Using Location Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing systems face challenges in efficiently classifying audio objects based on location metadata, particularly in real-time applications with many audio objects, leading to high computational complexity and potential audible artefacts.

Innovation Solution

A method that estimates the presence of dialog in audio objects using location metadata, assigning a confidence value to an object type parameter based on position, speed, acceleration, and elevation, allowing for selective dialog enhancement and clustering to reduce computational load and improve bitrate efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio objects are classified using traditional audio content analysis methods, then classification accuracy is improved, but computational complexity increases significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the classification task by using location metadata as a preliminary filter to identify candidate audio objects that may contain dialog, rather than analyzing all audio objects. This divides the computational workload into two stages: initial filtering based on spatial position and subsequent detailed analysis only for selected objects, thereby reducing overall computational complexity while maintaining classification accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces location metadata as an intermediary element that mediates between the audio content and the classification process. By using spatial position information as an intermediate filtering criterion, the system avoids direct complex analysis of all audio content, reducing computational load while preserving classification precision through the intermediate spatial selection step.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If dialog enhancement is applied to all audio objects, then dialog quality is improved, but bitrate efficiency deteriorates

Engineering Contradiction:
Improvedialog qualityVSAvoidbitrate efficiency
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent applies dialog enhancement selectively rather than uniformly across all audio objects. By using location metadata to identify which audio objects are likely to contain dialog (e.g., objects at specific spatial positions or with certain movement characteristics), the system applies enhancement only to those specific objects, thereby improving dialog quality where needed while avoiding unnecessary processing and maintaining bitrate efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter used for selecting audio objects for enhancement from traditional content-based features to location metadata parameters. This parameter change enables more efficient selection of candidates for dialog enhancement, improving the trade-off between dialog quality and bitrate efficiency by using spatial information as the selection criterion instead of computationally intensive content analysis.

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If clustering is performed on all audio objects, then audio scene organization is improved, but processing time increases

Engineering Contradiction:
Improveaudio scene organizationVSAvoidprocessing time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent performs preliminary filtering using location metadata before conducting clustering operations. By pre-selecting audio objects that are likely to contain dialog based on their spatial characteristics, the system reduces the number of objects that need to be clustered, thereby maintaining audio scene organization while significantly reducing processing time required for the clustering operation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11386913B2Audio object classification based on location metadata
Publication Date: 2022.07.12 DOLBY LABORATORIES LICENSING CORP
  • US11386913B2 patent drawing
  • US11386913B2 patent drawing
  • US11386913B2 patent drawing

AI summary

Methods (700, 800, 900), systems (200, 300, 400, 500, 600) and computer program products are provided. Location metadata (620) associated with an audio object is received (801). The location metadata defines a position of the audio object in an audio scene. It is estimated (630, 802), based on the location metadata, whether the audio object includes dialog. A value representative of a result of the estimation is assigned (803) to an object type parameter (231). In some example embodiments, audio objects are selected (661, 662, 804) based on values of their respective of object type parameters. In some example embodiments, at least one of the selected audio objects is submitted to dialog enhancement (690, 807).