Machine Learning Model Embedding Musical Features in Latent Space

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning systems face challenges in effectively learning and generating musical features that align with human perceptions, as they often struggle to balance feature learning for generative and retrieval-based applications, leading to suboptimal performance across tasks.

Innovation Solution

The proposed solution involves training a machine learning system with neural network components that embed musical features into a latent space with a predefined distribution, using techniques like adversarial training and backpropagation to optimize for context and distribution constraints, allowing for meaningful interpolation and sampling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a machine learning system is trained to learn musical features for generative applications, then the ability to generate musical content is improved, but the performance on retrieval-based applications deteriorates

Engineering Contradiction:
Improvegenerative application performanceVSAvoidretrieval-based application performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent trains a single machine learning system to perform multiple functions: both generative applications (creating new musical content) and retrieval-based applications (identifying and recommending existing music). The system uses a unified embedding space that can serve both purposes, eliminating the need for separate specialized models and enabling one system to handle diverse music-related tasks effectively

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If a machine learning system is trained to learn musical features for retrieval-based applications, then the ability to identify and retrieve musical content is improved, but the performance on generative applications deteriorates

Engineering Contradiction:
Improvemusical feature identification accuracyVSAvoidgenerative application capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system employs a universal embedding space that simultaneously supports precise musical feature representation for retrieval tasks and provides the structural foundation for generating new musical content. The same neural network components and loss functions are optimized to serve both retrieval accuracy and generative capability, allowing the model to excel at both identifying existing music and creating new compositions

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If multiple separate models are trained for different musical applications, then the performance of each specific application is improved, but the system complexity increases

Engineering Contradiction:
Improveapplication-specific performanceVSAvoidnumber of models and training processes
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple application-specific models into a single unified machine learning system. By combining the functionalities of separate generative and retrieval-based models into one system with shared embedding spaces and coordinated loss functions, the patent reduces overall system complexity while maintaining or improving performance across all applications

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11341945B2Techniques for learning effective musical features for generative and retrieval-based applications
Publication Date: 2022.05.24 SAMSUNG ELECTRONICS CO LTD
  • US11341945B2 patent drawing
  • US11341945B2 patent drawing
  • US11341945B2 patent drawing

AI summary

A method includes receiving a non-linguistic input associated with an input musical content. The method also includes, using a model that embeds multiple musical features describing different musical content and relationships between the different musical content in a latent space, identifying one or more embeddings based on the input musical content. The method further includes at least one of: (i) identifying stored musical content based on the one or more identified embeddings or (ii) generating derived musical content based on the one or more identified embeddings. In addition, the method includes presenting at least one of: the stored musical content or the derived musical content. The model is generated by training a machine learning system having one or more first neural network components and one or more second neural network components such that embeddings of the musical features in the latent space have a predefined distribution.