Automatic Interactive Audio Mashups With Vocal Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating music mashups requires specialized knowledge in music composition, making it difficult for most people to combine audio tracks effectively.

Innovation Solution

A system and method for combining audio tracks by separating vocal and accompaniment components, aligning structures, and adjusting tempo using a computer-implemented process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual mashup creation is performed by blending elements from multiple audio sources, then creative control and customization are improved, but the process becomes very difficult and requires specialized music composition knowledge

Engineering Contradiction:
Improvecreative controlVSAvoiddifficulty of creation
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system automatically segments audio tracks into distinct components (vocal, instrumental, beat, bass) using signal processing algorithms. This segmentation eliminates the need for users to manually separate audio elements, reducing the specialized knowledge required while maintaining creative control over which segments are combined.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary automated processing layer that handles complex audio analysis and manipulation tasks. This intermediary performs structure analysis, tempo detection, and component separation, bridging the gap between simple user input and complex audio manipulation outcomes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If automated audio processing is used to simplify mashup creation, then ease of operation is improved, but the ability to perform specialized audio manipulations may be reduced

Engineering Contradiction:
Improveease of creationVSAvoidaudio manipulation capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system provides dynamic control options that adapt to user needs. While automated processing handles basic tasks, users can intervene at multiple stages to adjust parameters, select specific components, and customize the blending process, maintaining versatility while improving ease of operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms that analyze the results of automated processing and allow users to refine outcomes. Users can review detected structures, adjust tempo alignments, and modify component selections, ensuring that automated processing enhances rather than limits creative capability.

Inventive Principle:
Principle #23Feedback

3Productivity

If structure analysis and alignment are performed automatically, then productivity is improved, but the complexity of the processing system increases

Engineering Contradiction:
Improvemashup creation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs self-service audio analysis by automatically detecting temporal structures, beats, and tempo without requiring external analysis tools. This self-contained approach improves productivity by eliminating manual analysis steps while managing complexity through integrated algorithms rather than external dependencies.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary structure analysis and tempo detection on all input tracks before the user begins the mashup process. This preliminary action prepares the data in advance, enabling faster user decision-making and reducing processing time during the actual creation phase, thereby improving overall productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12451106B2Automatic and interactive mashup system
Publication Date: 2025.10.21 LEMON INC(GB)
  • US12451106B2 patent drawing
  • US12451106B2 patent drawing
  • US12451106B2 patent drawing

AI summary

Systems and methods directed to combining audio tracks are provided. More specifically, a first audio track and a second audio track are received. The first audio track is separated into a vocal component and one or more accompaniment components. The second audio track is separated into a vocal component and one or more accompaniment components. A structure of the first audio track and a structure of the second audio track are determined. The first audio track and the second audio track are aligned based on the determined structures of the tracks. The vocal component of the first audio track is stretched to match a tempo of the second audio track. The stretched vocal component of the first audio track is added to the one or more accompaniment components of the second audio track.