Audio Object Selection via Reference-Guided Decomposition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies lack effective methods for users to select and manipulate specific sound objects within audio mixtures, such as isolating a singer's voice from a musical recording, due to the complexity of visualizing and interacting with mixed sounds.
Innovation Solution
A system and method that allows users to select a target sound object from a mixture by providing reference audio data, using Probabilistic Latent Component Analysis (PLCA) and iterative Expectation-Maximization (EM) algorithms to decompose and re-synthesize the audio data, enabling object-based interaction and manipulation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Illumination intensity
If waveform representation is used to visualize audio data, then accurate visualization of sound pressure over time is achieved, but the amount of useful information for identifying sound objects remains very limited
Solution Approach 1:
The patent transforms the audio data from a one-dimensional waveform representation (time only) into a two-dimensional time-frequency representation (spectrogram). This dimensional transformation reveals frequency information that is invisible in the waveform, enabling users to distinguish between different sound objects based on their spectral characteristics while maintaining temporal context.
2Loss of information
If time-frequency visualizations (spectrograms) are used to show frequency energy content, then visual distinction between mixed sounds is improved, but object-based interaction and selection capability is lost
Solution Approach 1:
The patent introduces an intermediary processing layer that takes the spectrogram as input and applies reference-based signal separation algorithms. This intermediary layer translates the visual frequency information into actionable audio object separation, allowing users to interact with specific sound objects by providing reference examples, thus bridging the gap between visualization and manipulation.
Solution Approach 2:
The patent uses reference audio samples provided by the user as templates or copies of the target sound object. By comparing the spectrogram features against these reference copies, the system identifies and isolates matching sound objects in the mixture, enabling intuitive object-based interaction through simple reference recording.
3Device complexity
If traditional audio processing methods are used, then computational simplicity is maintained, but the ability to select and manipulate specific sound objects from mixtures is lost
Solution Approach 1:
The patent segments the mixed audio signal into distinct sound objects by analyzing the spectrogram and separating frequency-temporal regions corresponding to different sources. This segmentation approach breaks down the complex mixture into manageable, independently controllable components, enabling selective manipulation of individual sound objects while maintaining reasonable computational complexity.
Data Source
AI summary
A system and method are described for selecting a target sound object from a sound mixture. In embodiments, a sound mixture comprises a plurality of sound objects superimposed in time. A user can select one of these sound objects by providing reference audio data corresponding to a reference sound object. The system analyzes the audio data and the reference audio data to identify a portion of the audio data corresponding to a target sound object in the mixture that is most similar to the reference sound object. The analysis may include decomposing the reference audio data into a plurality of reference components and the sound mixture into a plurality of components guided by the reference components. The target sound object can be re-synthesized from the target components.


