Generative Instrument Tone Addition Using Genre-Aware Audio Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for adding musical instrument tones to recorded songs are labor-intensive, time-consuming, and prone to errors, lacking the ability to intuitively adapt to the genre and contextual awareness of the song.

Innovation Solution

An audio synthesis model combining WaveNet and Generative Adversarial Network (GAN) is trained on an audio dataset to generate musical instrument tones in conformity with the genre of a vocal track, using spectrogram processing and refined temporal sequencing to create a cohesive audio track through a Simple Additive Mixing operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional manual recording and mixing methods are used to add musical instrument tones, then the process allows for human creativity and control, but it becomes labor-intensive, time-consuming, and error-prone

Engineering Contradiction:
Improveautomation of instrument tone additionVSAvoidcomplexity of audio synthesis system
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical recording and mixing operations with an automated audio synthesis system using deep learning models (WaveNet and GAN). The system automatically generates instrument tones by processing vocal tracks through neural networks, eliminating the need for manual recording sessions and mixing operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The audio synthesis system performs self-service by automatically analyzing the input vocal track, determining appropriate instrument tones based on genre and temporal dependencies, generating the tones through the synthesis model, and mixing them with the original track without requiring human intervention at each step.

Inventive Principle:
Principle #25Self-service

2Loss of time

If manual re-recording is performed to incorporate new instrument tones, then the song quality can be maintained, but the process becomes time-consuming and requires access to original recording sessions

Engineering Contradiction:
Improvetime required to add instrument tonesVSAvoidaccuracy of genre and temporal alignment
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system performs preliminary analysis of the vocal track to extract temporal dependencies and genre characteristics before generating instrument tones. The audio synthesis model is pre-trained on extensive music data to understand genre conventions and temporal patterns, enabling accurate tone generation without time-consuming manual adjustments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback mechanisms where the generated instrument tones are evaluated against the original vocal track's temporal structure and genre characteristics. The synthesis model adjusts its output based on feedback from the temporal dependency analysis to ensure accurate alignment with the song's rhythm and style.

Inventive Principle:
Principle #23Feedback

3Productivity

If automated audio synthesis is used to generate instrument tones, then real-time processing is achieved, but ensuring genre conformity and temporal accuracy becomes challenging

Engineering Contradiction:
Improvespeed of adding instrument tonesVSAvoidprecision of temporal sequencing and genre alignment
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The audio synthesis process is segmented into distinct functional modules: temporal dependency extraction from the vocal track, genre determination based on audio features, instrument tone generation through the synthesis model, and final mixing. This segmentation allows each module to specialize in one aspect, improving both speed and precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically changes parameters such as tempo, key, and instrumentation based on the analyzed characteristics of the input vocal track. The audio synthesis model adjusts these parameters in real-time to ensure the generated instrument tones conform to the detected genre and maintain accurate temporal alignment with the original recording.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250259611A1Generative addition of musical instrument tones to songs
Publication Date: 2025.08.14 INFOSYS LTD
  • US20250259611A1 patent drawing
  • US20250259611A1 patent drawing
  • US20250259611A1 patent drawing

AI summary

Generative filling of musical instrument tones to songs is provided. For a vocal track, a genre and musical instruments to be added to the vocal track are determined. Using an audio synthesis model, an audio tone is generated for each musical instrument in conformity with the genre. Further, each audio tone is converted into a spectrogram, which when processed based on temporal dependencies, generates a refined temporal sequence for the audio tone. Based on the refined temporal sequence of each audio tone, an audio waveform is generated. A simple additive mixing operation is executed on the vocal track and the audio waveforms generated for the musical instruments to generate an audio track (e.g., a new song).