Audio Stem Identification Using Lower-Dimensional Vector Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems lack an effective method to identify which audio stems in a database have been used to create a music track, especially as the database grows, as users cannot listen to all stems to make their choices.

Innovation Solution

A machine learning-based approach is employed to train a model that predicts the probability of audio stems being related to input stems, using collaborative filtering and audio similarity algorithms, and maps acoustic feature vectors to a lower-dimensional vector space to identify complementary stems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users manually listen to all stems in the database to identify which stems have been used to make a music track, then the identification accuracy is high, but the time consumption and operational complexity increase significantly as the database grows

Engineering Contradiction:
Improvestem identification accuracyVSAvoidtime to identify stems
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of listening to and evaluating audio stems with an automated audio analysis system. The system uses audio fingerprinting, spectral analysis, and machine learning algorithms to automatically identify stems that have been used in a music track, eliminating the need for manual listening while maintaining high identification accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary audio analysis system that acts as a mediator between the user's identification need and the stem database. This system processes audio signals, extracts features, compares them against the database, and returns identification results, thereby reducing the time and effort required for stem identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the stem database grows larger to provide more options for users, then the versatility and choice increase, but the complexity of identifying which stems have been used increases

Engineering Contradiction:
Improvestem selection flexibilityVSAvoididentification system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the identification process into distinct modular components: audio preprocessing, feature extraction, database searching, and result generation. Each module handles a specific aspect of the identification task, making the overall system more manageable and scalable as the database grows. This modular architecture allows the system to efficiently handle larger databases without proportionally increasing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates compressed representations (audio fingerprints) of the actual audio stems that serve as copies for comparison purposes. These fingerprints capture the essential characteristics of each stem in a compact form, allowing efficient database searching and comparison without requiring the system to process the full-resolution audio data for every stem in the database.

Inventive Principle:
Principle #26Copying

3Ease of operation

If existing tag-based search systems are used to identify stems, then the ease of operation is maintained, but the identification accuracy and comprehensiveness decrease

Engineering Contradiction:
Improvesearch system usabilityVSAvoidstem identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent merges multiple identification approaches into a unified system that combines audio-based identification with tag-based searching. The system first performs automated audio analysis to identify stems with high precision, then integrates tag-based filtering to maintain ease of operation. This combination allows users to benefit from both the accuracy of audio analysis and the simplicity of tag-based interfaces.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250218453A1Audio stem identification systems and methods
Publication Date: 2025.07.03 SPOTIFY
  • US20250218453A1 patent drawing
  • US20250218453A1 patent drawing
  • US20250218453A1 patent drawing

AI summary

Methods, systems and computer program products are provided for determining acoustic feature vectors of query and target items in a first vector space, and mapping the acoustic feature vectors to a second vector space having a lower dimension. The distribution of vectors in the second vector space can then be used to identify items from the same songs, and/or items that are complementary. A mapping function is trained using a machine learning algorithm, such that complementary audio items are closer in the second vector space than the first, according to a given distance metric.