Latent-Space Audio Embeddings for Psychoacoustic Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio content retrieval systems struggle with subjective and multi-factored attributes of digital audio signals, such as timbre, rhythm, and melody, as they rely on keyword-based tagging and manual feature selection, which often omit useful features and use redundant ones, making it difficult to find similar sounds.

Innovation Solution

A system that uses machine learning techniques to generate contextual latent-space representations of digital audio signals through an artificial neural network, learning unsupervised sound embeddings that encode psychoacoustic attributes, allowing for more accurate similarity detection and recommendation without manual labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword-based tagging and manual feature selection are used for audio content retrieval, then the system is simple to implement and operate, but it omits useful discriminative features and uses redundant features, leading to poor similarity detection accuracy

Engineering Contradiction:
Improvesimilarity detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual feature selection and keyword-based retrieval (mechanical/manual system) with an automated neural network-based latent space representation system. The neural network automatically learns and extracts discriminative features from audio signals, substituting the manual tagging process with an automated machine learning approach that improves accuracy while managing complexity through algorithmic automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network performs self-service by automatically learning feature representations from raw audio data without requiring manual feature engineering or keyword tagging. The system autonomously identifies discriminative features and constructs latent space representations, eliminating the need for human experts to manually select and tag features, thereby improving accuracy while the automation manages the complexity.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual feature selection methods are used, then full knowledge and control over signal representation is achieved, but useful discriminative features are often omitted and redundant features are used

Engineering Contradiction:
Improvefeature selection reliabilityVSAvoidfeature discrimination precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent substitutes manual feature selection with an automated neural network-based feature learning system. The neural network automatically identifies and extracts discriminative features from audio signals, replacing the manual process that suffers from human limitations in identifying all useful features. This automated approach improves feature discrimination precision while maintaining reliability through consistent algorithmic application.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the feature representation parameters by learning latent space representations that capture essential audio characteristics. Instead of using fixed manual features, the system learns optimal feature parameters through training, transforming the feature space to maximize discrimination between different audio contents while ensuring reliable and consistent feature extraction.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If keyword-based search is used for audio retrieval, then objective attributes can be searched effectively, but subjective and multi-factored attributes such as timbre, rhythm, and melody cannot be effectively searched

Engineering Contradiction:
Improvesearch attribute versatilityVSAvoidsubjective attribute search precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms audio attributes into a unified latent space representation that captures both objective and subjective characteristics. By encoding timbre, rhythm, melody, and other subjective attributes into numerical vectors in latent space, the system enables mathematical operations and similarity measurements that were not possible with traditional keyword-based approaches, thereby improving search precision for subjective attributes while maintaining versatility.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transitions from one-dimensional keyword matching to multi-dimensional latent space representations. Each audio signal is represented by a vector in latent space that captures multiple attributes simultaneously, enabling the system to search and compare audio based on complex combinations of subjective attributes like timbre, rhythm, and melody, thereby improving search versatility and precision for subjective characteristics.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Productivity

If unsupervised learning is used to generate sound embeddings, then manual labeling is eliminated and processing speed increases, but the quality of sound embeddings must be ensured without supervision

Engineering Contradiction:
Improveprocessing speedVSAvoidembedding quality precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements self-service through unsupervised learning where the neural network automatically learns to generate meaningful sound embeddings without manual labeling. The system autonomously identifies patterns and structures in audio data, eliminating the need for time-consuming manual annotation while maintaining embedding quality through the network's ability to learn discriminative features directly from the data, thereby improving processing speed without sacrificing precision.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12051439B2Method and system for learning and using latent-space representations of audio signals for audio content-based retrieval
Publication Date: 2024.07.30 DISTRIBUTED CREATION INC
  • US12051439B2 patent drawing
  • US12051439B2 patent drawing
  • US12051439B2 patent drawing

AI summary

A method and system are provided for extracting features from digital audio signals which exhibit variations in pitch, timbre, decay, reverberation, and other psychoacoustic attributes and learning, from the extracted features, an artificial neural network model for generating contextual latent-space representations of digital audio signals. A method and system are also provided for learning an artificial neural network model for generating consistent latent-space representations of digital audio signals in which the generated latent-space representations are comparable for the purposes of determining psychoacoustic similarity between digital audio signals. A method and system are also provided for extracting features from digital audio signals and learning, from the extracted features, an artificial neural network model for generating latent-space representations of digital audio signals which take care of selecting salient attributes of the signals that represent psychoacoustic differences between the signals.