Sound Mixture Recognition via Spectral Basis Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in efficiently processing and searching multimedia content, particularly in separating and identifying multiple sound sources within audio mixtures, as computers struggle to differentiate between overlapping sound sources like music, dialog, and noise.
Innovation Solution
The proposed solution involves a method to estimate the proportions of sound sources in a mixture using a composite model that includes spectral basis vectors and a transition matrix, which represents temporal dependencies, allowing for iterative weight refinement without actual source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio processing methods are used to analyze sound mixtures, then the processing is simpler, but the ability to differentiate between constituent sound sources is insufficient
Solution Approach 1:
The sound mixture is segmented into multiple sound sources by representing each source with its own spectral basis vectors and transition matrices. The composite model divides the overall sound analysis into separate source-specific components, allowing the system to differentiate between multiple concurrent sound sources in the mixture.
Solution Approach 2:
The system changes parameters by introducing spectral basis vectors and transition matrices to characterize each sound source. By varying the spectral representation and temporal dynamics parameters across different sources, the model achieves better differentiation accuracy while managing complexity through parameterized source representations.
2Measurement precision
If source separation techniques are applied to separate sound sources, then the source identification improves, but the computational complexity increases
Solution Approach 1:
The patent extracts source identification information directly from the sound mixture without performing full source separation. By taking out only the necessary spectral and temporal characteristics needed for identification, the system achieves accurate source differentiation while avoiding the computational burden of complete source separation.
Solution Approach 2:
Instead of performing complete source separation, the system applies partial action by estimating source proportions and characteristics using spectral basis vectors and transition matrices. This partial approach provides sufficient identification accuracy for many applications without the excessive computational cost of full separation.
3Measurement precision
If temporal information is incorporated into the model, then the source recognition accuracy improves, but the model complexity increases
Solution Approach 1:
The transition matrices are pre-computed to capture temporal dependencies between spectral basis vectors. By performing this temporal analysis in advance and storing it in matrix form, the model incorporates temporal information for improved recognition accuracy while avoiding the need for complex real-time temporal computations during source identification.
Data Source
AI summary
A sound mixture may be received that includes a plurality of sources. A model may be received that includes a dictionary of spectral basis vectors for the plurality of sources. A weight may be estimated for each of the plurality of sources in the sound mixture based on the model. In some examples, such weight estimation may be performed using a source separation technique without actually separating the sources.


