Audio Object Extraction via Spectral Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies lack an effective method for accurately and efficiently extracting audio objects from traditional channel-based audio content, which limits the ability to provide immersive experiences similar to object-based audio formats.

Innovation Solution

A method and system for audio object extraction that applies frame-level audio object extraction based on frequency spectral similarities among channels, followed by audio object composition across frames to generate complete tracks of audio objects, utilizing techniques such as hierarchical clustering and probability matrix generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If frame-level audio object extraction is applied to extract audio objects from channel-based content, then the ability to provide immersive experiences is improved, but the complexity of the extraction process increases

Engineering Contradiction:
Improveability to provide immersive experiencesVSAvoidcomplexity of extraction process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The extraction process is divided into two distinct stages: frame-level audio object extraction and audio object composition across frames. This segmentation allows the system to handle complex extraction tasks in manageable steps, first identifying objects within individual frames using frequency spectral similarities, then composing these objects across multiple frames to maintain temporal consistency and provide immersive experiences.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary audio object extraction on individual frames before composing objects across frames. By pre-identifying audio objects in each frame based on frequency spectral similarities among channels, the system prepares the necessary data structures and object identifiers that will be used in the subsequent composition stage, reducing the overall computational complexity.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If audio object extraction is performed on individual frames based on frequency spectral similarities, then the accuracy of object identification is improved, but the computational time increases

Engineering Contradiction:
Improveaccuracy of object identificationVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies audio object extraction only on individual frames where it is necessary to identify audio objects, rather than processing every frame with the full extraction algorithm. By selectively applying extraction based on frequency spectral similarities, the system achieves accurate object identification while reducing unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system extracts only the essential information needed for audio object identification from each frame, specifically focusing on frequency spectral similarities among channels. By taking out only the critical features required for object identification rather than processing all audio data, the system improves accuracy while minimizing computational time requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If audio object composition is performed across frames, then the completeness of audio object tracks is improved, but the processing complexity increases

Engineering Contradiction:
Improvecompleteness of audio object tracksVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges audio object information across multiple frames by composing extracted objects temporally. By combining the results from individual frame extractions and maintaining temporal consistency through composition operations, the system produces complete and reliable audio object tracks that span across the entire audio content duration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The audio object composition process maintains continuous tracking of audio objects across frames, ensuring that objects are followed and tracked throughout the temporal sequence. This continuity ensures complete audio object tracks are generated, with the system maintaining useful action by consistently applying composition operations across all relevant frames.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9786288B2Audio object extraction
Publication Date: 2017.10.10 DOLBY LABORATORIES LICENSING CORP
  • US9786288B2 patent drawing
  • US9786288B2 patent drawing
  • US9786288B2 patent drawing

AI summary

Embodiments of the present invention relate to audio object extraction. A method for audio object extraction from audio content of a format based on a plurality of channels is disclosed. The method comprises applying audio object extraction on individual frames of the audio content at least partially based on frequency spectral similarities among the plurality of channels. The method further comprises performing audio object composition across the frames of the audio content, based on the audio object extraction on the individual frames, to generate a track of at least one audio object. Corresponding system and computer program product are also disclosed.