Music File Processing via Vocal-Accompaniment Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing music processing technologies lack flexibility in sound effect adjustments, leading to unsuitable sound effects for music files, resulting in a poor user experience due to fixed sound effect types that do not adapt to the specific characteristics of the music.

Innovation Solution

A method and device that extract voice and accompaniment data from a music file, apply separate sound effect processing to each, and synthesize them based on parameters like rhythm, motion trajectory, and accompaniment mode, allowing for targeted adjustments in pitch, timbre, loudness, and dynamic range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If fixed sound effect types are applied to the entire music file, then the processing is simple and uniform, but the flexibility and adaptability to different music characteristics are poor

Engineering Contradiction:
Improveadaptability of sound effect to music characteristicsVSAvoidcomplexity of sound effect processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The music file is segmented into different audio components (voice, accompaniment, and specific musical instruments) using audio source separation technology. This allows different sound effect processing to be applied to each component independently, improving adaptability while managing complexity through modular processing of separated audio signals

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different sound effect parameters are applied to different audio components based on their specific characteristics. For example, pitch adjustment is applied to voice data while playing orientation adjustment is applied to accompaniment data, allowing each component to receive locally optimized sound effects rather than uniform processing

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If separate sound effect processing is applied to voice and accompaniment data, then the flexibility and play effect are improved, but the processing complexity increases

Engineering Contradiction:
Improveprecision of sound effect adjustmentVSAvoidcomplexity of audio processing device
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The audio signal is segmented into voice data and accompaniment data through audio source separation, enabling precise independent processing of each component. This segmentation allows targeted sound effect adjustments (pitch for voice, playing orientation for accompaniment) without requiring complex full-signal processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The audio processing device integrates multiple functions including audio source separation, pitch adjustment, playing orientation adjustment, and synthesis into a single system. This multi-functionality achieves precise sound effect control while consolidating complexity into a unified processing architecture

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3940690B1Method and device for processing music file, terminal and storage medium
Publication Date: 2024.11.06 DOUYIN VISION CO LTD
  • EP3940690B1 patent drawingFigure 1
  • EP3940690B1 patent drawingFigure 2
  • EP3940690B1 patent drawingFigure 3

AI summary

Provided are a method and device for processing a music file, a terminal and a storage medium. The method comprises: in response to a received sound effect adjustment instruction, acquiring a music file, the adjustment of which is indicated by the sound effect adjustment instruction; carrying out vocals and accompaniment separation on the music file to obtain vocal data and accompaniment data in the music file; carrying out first sound effect processing on the vocal data to obtain target vocal data, and carrying out second sound effect processing on the accompaniment data to obtain target accompaniment data; and synthesizing the target vocal data and the target accompaniment data to obtain a target music file.