Variational Autoencoder for Audio Style Vector Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in navigating large media collections to find content of interest, as existing systems require significant user input and do not effectively convey musical style information, making it difficult to recommend or stream relevant audio content.

Innovation Solution

A method using a variational autoencoder (VAE) to generate representative vectors for audio content items, which are then used to create a vector space representing musical style similarity, allowing for efficient content recommendation and selection based on user preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users navigate through large content collections manually to find content of interest, then users can access content items, but the process becomes challenging and time-consuming

Engineering Contradiction:
Improveease of content navigationVSAvoidtime to find content
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual navigation (mechanical user action) with an automated recommendation system that uses machine learning models to automatically analyze and recommend content based on user preferences and listening history, eliminating the need for manual browsing through large collections

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically generating content recommendations without requiring active user input for each recommendation, using pre-collected user preference data and listening history to autonomously curate content suggestions

Inventive Principle:
Principle #25Self-service

2Loss of information

If the platform provides basic information about content items (e.g., song title), then information is available, but the information is insufficient to help users decide whether to play back the content

Engineering Contradiction:
Improveinformation sufficiency for content selectionVSAvoidsystem complexity for generating recommendations
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent adds a new dimension of information by transforming audio content into vector representations in a multi-dimensional vector space, where geometric relationships encode musical style similarities, providing rich contextual information beyond basic metadata like song titles

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system performs preliminary action by pre-processing and encoding audio content into vector representations in advance, storing these in a vector space index that enables rapid similarity search and recommendation generation without requiring complex processing at the moment of user query

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If representative vectors are generated for multiple audio content items to create a vector space, then musical style similarity can be determined, but processing complexity increases

Engineering Contradiction:
Improveprecision of musical style similarity measurementVSAvoidcomplexity of vector processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes parameters by transforming audio data into a different parameter space (vector space) where similarity measurements become simpler geometric operations, enabling precise musical style comparison through distance calculations rather than complex audio analysis

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The vector space acts as an intermediary representation between raw audio content and similarity measurement, providing a simplified medium where complex audio style comparisons are reduced to straightforward geometric distance calculations

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11887613B2Determining musical style using a variational autoencoder
Publication Date: 2024.01.30 SPOTIFY
  • US11887613B2 patent drawing
  • US11887613B2 patent drawing
  • US11887613B2 patent drawing

AI summary

A computer extracts a vocal portion from a first audio content item and determines a first representative vector that corresponds to a vocal style of the first audio content item by applying a variational autoencoder (VAE) to the extracted vocal portion of the representation of the audio content item. The computer streams, to an electronic device, a second audio content item, selected from a plurality of audio content items, that has a second representative vector that corresponds to a vocal style of the second audio content item, wherein the second representative vector that corresponds the vocal style of the second audio content item meets similarity criteria with respect to the first representative vector that corresponds to the vocal style of the first audio content item.