ML Source Separation for Music Stem Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current music visualization techniques, such as waveform representations, provide limited information for DJs, requiring pre-preparation of cue points and lacking real-time mixing capabilities, especially when handling multiple songs.
Innovation Solution
The use of machine learning source separation models to generate compact visual representations of music stems, including fundamental frequencies and amplitudes, allowing for improved on-the-fly mixing by providing detailed information within a structured graphical user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional waveform representations are used for music visualization, then the visualization is simple and easy to generate, but the information provided is limited and lacks detail for real-time mixing
Solution Approach 1:
The music stem data is segmented into multiple visual channels including amplitude channel, frequency channel, and highlight channel. Each channel processes and visualizes specific aspects of the audio data independently, allowing comprehensive information display without overwhelming complexity in a single visualization stream
Solution Approach 2:
The patent adds vertical positioning dimension to represent fundamental frequency information, while horizontal position represents time and vertical extent represents amplitude. This multi-dimensional visualization approach provides rich audio information (frequency, amplitude, timing) without increasing interface complexity, as all dimensions are integrated into a single spectrogram-style display
2Ease of operation
If detailed music analysis is performed to provide comprehensive mixing information, then the mixing capability is improved, but the preparation time and processing complexity increase
Solution Approach 1:
The system performs preliminary analysis of the music stem data to extract fundamental frequencies, amplitudes, and temporal information before the mixing operation begins. This pre-processing creates a ready-to-use visual representation that enables immediate real-time mixing without requiring additional preparation time during actual use
Solution Approach 2:
The patent replaces manual audio analysis and cue point preparation with automated machine learning source separation models and signal processing algorithms. The system automatically generates the visual representation and extracts mixing information, eliminating the need for manual preparation work while providing comprehensive mixing capabilities
Data Source
AI summary
Devices, methods, and other aspects are provided for compact visual representation of song data. A device may receive an indication of a first song selection from a song data listing, receive processed stem data associated with the first song selection, wherein the processed stem data is generated from first song music data processed by a machine learning source separation model, and dynamically generate first song frequency indicators and first song amplitude indicators from the processed stem data for times from a start to an end of the first song music data. The device may dynamically display the first song amplitude indicators in a timing window of a user interface in various ways in accordance with aspects described.


