Generative Audio Models for Vocal Accompaniment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating instrumental audio to accompany vocals are labor-intensive and require significant expertise or resources, making it difficult for users to produce high-quality musical accompaniments efficiently.

Innovation Solution

A computer-implemented method using machine-learned generative audio models processes vocal and instrumental audio data to generate predicted instrumental audio, trained on separated data from musical works, allowing for efficient creation of musical accompaniments by minimizing the difference between vocal and instrumental intermediate representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to generate instrumental audio to accompany vocals, then high-quality musical accompaniments can be produced, but the process becomes labor-intensive and requires significant expertise or resources

Engineering Contradiction:
Improvequality of musical accompanimentVSAvoidcomplexity of generation process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual, expert-driven methods for generating instrumental accompaniments with an automated machine learning system. The neural network model processes vocal audio input and generates corresponding instrumental audio automatically, substituting the need for human musicians or complex audio production workflows with an computational algorithm that learns from training data consisting of paired vocal and instrumental audio segments.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If conventional methods are used to generate instrumental audio to accompany vocals, then high-quality musical accompaniments can be produced, but significant expertise or resources are required

Engineering Contradiction:
Improvequality of musical accompanimentVSAvoidease of generating accompaniment
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system enables users to generate instrumental accompaniments independently without requiring expertise in music production or audio engineering. The machine learning model handles the complex task of creating harmonious instrumental arrangements that match the vocal input, allowing users with basic audio recording capabilities to produce professional-quality musical accompaniments through a simplified interface that automatically processes the vocal recording and generates the corresponding instrumental track.

Inventive Principle:
Principle #25Self-service

3Productivity

If machine learning models are used to generate instrumental audio from vocal input, then the generation process becomes efficient and accessible, but computational resources are required

Engineering Contradiction:
Improvespeed of generating accompanimentVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary training using large datasets of paired vocal and instrumental audio during an offline phase, allowing the model to learn the complex relationships between vocals and accompaniments in advance. During actual use, the pre-trained model can rapidly generate instrumental audio from new vocal inputs with reduced computational requirements, as the heavy lifting of learning the generation process has already been completed during the initial training phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240395233A1Machine-Learned Models for Generation of Musical Accompaniments Based on Input Vocals
Publication Date: 2024.11.28 GOOGLE LLC
  • US20240395233A1 patent drawing
  • US20240395233A1 patent drawing
  • US20240395233A1 patent drawing

AI summary

Training data comprising a plurality of training pairs is obtained. Each training pair comprises instrumental audio data and vocal audio data separated from audio data of a musical work of a respective plurality of musical works. For one or more training pairs of the plurality of training pairs, the vocal audio data is processed with machine-learned model(s) of a machine-learned generative audio model grouping to obtain a vocal intermediate representation for the vocal audio data. The instrumental audio data is processed with a pre-trained encoding model to obtain an instrumental intermediate representation for the instrumental audio data. A loss function is evaluated that evaluates a difference between the vocal intermediate representation and the instrumental intermediate representation. Values of parameters of a machine-learned model of the machine-learned generative audio model grouping are modified based on the loss function.