Conditional Similarity Networks for Cross-Format Music Relations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to determine relations between music items of different formats and lengths, such as audio and symbolic files, limiting the ability to identify related tracks and submixes within large databases.
Innovation Solution
The use of conditional similarity networks (CSNs) to map music items into multiple subspaces, each modeling different characteristics, enabling cross-domain relation and diverse ways of relating music files, including complementarity, consecutiveness, and mood similarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing similarity methods are applied to music files, then similarity between whole songs can be estimated, but relations between individual tracks and submixes of different formats cannot be determined
Solution Approach 1:
The patent segments music files into individual tracks and submixes, applying source separation techniques to decompose whole songs into constituent elements. This allows the system to analyze and relate individual musical components (drums, bass, vocals, etc.) rather than treating entire songs as monolithic units, thereby enabling format-agnostic comparison across audio and symbolic representations.
Solution Approach 2:
The patent introduces intermediate representations (feature vectors, embeddings) that serve as mediators between different music file formats. By converting audio files, MIDI files, and symbolic representations into common intermediate feature spaces, the system enables reliable comparison and relation determination across diverse formats without direct format-to-format mapping.
2Adaptability or versatility
If CSN techniques are extended to individual tracks, then diverse music files can be related, but the complexity of the system increases
Solution Approach 1:
The patent develops universal models that perform multiple functions: they can process both audio and symbolic representations, handle individual tracks and submixes, and determine various types of relations (similarity, complementarity, consecutiveness). This multi-functional approach increases versatility while managing system complexity through shared architectural components and unified processing pipelines.
Solution Approach 2:
The patent employs parameter changes by adjusting model configurations and processing parameters based on the specific task and input type. Different relation types (similarity vs. complementarity) are handled by modifying model parameters and loss functions, allowing the same base system to adapt to diverse requirements without requiring completely separate architectures for each function.
3Adaptability or versatility
If music files of different lengths are processed, then flexibility is improved, but processing time and computational resources increase
Solution Approach 1:
The patent implements dynamic processing that adapts to varying input lengths. The models can process anything from individual bars to full-length songs by dynamically adjusting the temporal scope of analysis. This is achieved through flexible sequence processing architectures that can handle variable-length inputs without requiring fixed-time processing windows, optimizing computational resources based on actual input requirements.
Data Source
AI summary
A method of determining relations between music items, the method comprising determining a first input representation for a symbolic representation of a first music item, mapping the first input representation onto to one or more subspaces derived from a vector space using a first model, wherein each subspace models a characteristic of the music items, determining a second input representation for music data representing a second music item, mapping the second input representation onto the one or more subspaces using a second model, determining a distance between the mappings of the first and second input representation in each subspace, wherein the distance represents the degree of relation between the first and second input representation with respect to the characteristic modelled by the subspace.


