ML Source Separation for Music Stem Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current music visualization techniques, such as waveform representations, provide limited information for DJs, requiring pre-preparation of cue points and lacking real-time mixing capabilities, especially when handling multiple songs.

Innovation Solution

The use of machine learning source separation models to generate compact visual representations of music stems, including fundamental frequencies and amplitudes, allowing for improved on-the-fly mixing by providing detailed information within a structured graphical user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional waveform representations are used for music visualization, then the visualization is simple and easy to generate, but the information provided is limited and lacks detail for real-time mixing

Engineering Contradiction:
Improveinformation detailVSAvoidvisualization complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The music stem data is segmented into multiple visual channels including amplitude channel, frequency channel, and highlight channel. Each channel processes and visualizes specific aspects of the audio data independently, allowing comprehensive information display without overwhelming complexity in a single visualization stream

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds vertical positioning dimension to represent fundamental frequency information, while horizontal position represents time and vertical extent represents amplitude. This multi-dimensional visualization approach provides rich audio information (frequency, amplitude, timing) without increasing interface complexity, as all dimensions are integrated into a single spectrogram-style display

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If detailed music analysis is performed to provide comprehensive mixing information, then the mixing capability is improved, but the preparation time and processing complexity increase

Engineering Contradiction:
Improvemixing capabilityVSAvoidpreparation time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the music stem data to extract fundamental frequencies, amplitudes, and temporal information before the mixing operation begins. This pre-processing creates a ready-to-use visual representation that enables immediate real-time mixing without requiring additional preparation time during actual use

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual audio analysis and cue point preparation with automated machine learning source separation models and signal processing algorithms. The system automatically generates the visual representation and extracts mixing information, eliminating the need for manual preparation work while providing comprehensive mixing capabilities

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240312441A1Machine learning assisted music visualization
Publication Date: 2024.09.19 HILSS ADAM
  • US20240312441A1 patent drawing
  • US20240312441A1 patent drawing
  • US20240312441A1 patent drawing

AI summary

Devices, methods, and other aspects are provided for compact visual representation of song data. A device may receive an indication of a first song selection from a song data listing, receive processed stem data associated with the first song selection, wherein the processed stem data is generated from first song music data processed by a machine learning source separation model, and dynamically generate first song frequency indicators and first song amplitude indicators from the processed stem data for times from a start to an end of the first song music data. The device may dynamically display the first song amplitude indicators in a timing window of a user interface in various ways in accordance with aspects described.