Audio Object Selection via Reference-Guided Decomposition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing technologies lack effective methods for users to select and manipulate specific sound objects within audio mixtures, such as isolating a singer's voice from a musical recording, due to the complexity of visualizing and interacting with mixed sounds.

Innovation Solution

A system and method that allows users to select a target sound object from a mixture by providing reference audio data, using Probabilistic Latent Component Analysis (PLCA) and iterative Expectation-Maximization (EM) algorithms to decompose and re-synthesize the audio data, enabling object-based interaction and manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If waveform representation is used to visualize audio data, then accurate visualization of sound pressure over time is achieved, but the amount of useful information for identifying sound objects remains very limited

Engineering Contradiction:
Improvevisualization accuracyVSAvoidinformation content
Core Design Contradiction:
Illumination intensityVSLoss of information

Solution Approach 1:

The patent transforms the audio data from a one-dimensional waveform representation (time only) into a two-dimensional time-frequency representation (spectrogram). This dimensional transformation reveals frequency information that is invisible in the waveform, enabling users to distinguish between different sound objects based on their spectral characteristics while maintaining temporal context.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If time-frequency visualizations (spectrograms) are used to show frequency energy content, then visual distinction between mixed sounds is improved, but object-based interaction and selection capability is lost

Engineering Contradiction:
Improveinformation contentVSAvoidobject-based interaction
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent introduces an intermediary processing layer that takes the spectrogram as input and applies reference-based signal separation algorithms. This intermediary layer translates the visual frequency information into actionable audio object separation, allowing users to interact with specific sound objects by providing reference examples, thus bridging the gap between visualization and manipulation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses reference audio samples provided by the user as templates or copies of the target sound object. By comparing the spectrogram features against these reference copies, the system identifies and isolates matching sound objects in the mixture, enabling intuitive object-based interaction through simple reference recording.

Inventive Principle:
Principle #26Copying

3Device complexity

If traditional audio processing methods are used, then computational simplicity is maintained, but the ability to select and manipulate specific sound objects from mixtures is lost

Engineering Contradiction:
Improveprocessing complexityVSAvoidsound object selection capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the mixed audio signal into distinct sound objects by analyzing the spectrogram and separating frequency-temporal regions corresponding to different sources. This segmentation approach breaks down the complex mixture into manageable, independently controllable components, enabling selective manipulation of individual sound objects while maintaining reasonable computational complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8954175B2User-guided audio selection from complex sound mixtures
Publication Date: 2015.02.10 ADOBE INC
  • US8954175B2 patent drawing
  • US8954175B2 patent drawing
  • US8954175B2 patent drawing

AI summary

A system and method are described for selecting a target sound object from a sound mixture. In embodiments, a sound mixture comprises a plurality of sound objects superimposed in time. A user can select one of these sound objects by providing reference audio data corresponding to a reference sound object. The system analyzes the audio data and the reference audio data to identify a portion of the audio data corresponding to a target sound object in the mixture that is most similar to the reference sound object. The analysis may include decomposing the reference audio data into a plurality of reference components and the sound mixture into a plurality of components guided by the reference components. The target sound object can be re-synthesized from the target components.