Dual-Source Audio Synthesis via Energy Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies lack an efficient method to automatically synthesize music from two songs with the same accompaniment but different vocals, requiring manual selection and processing, which is cumbersome and resource-intensive.

Innovation Solution

A dual sound source audio data processing method that identifies and filters songs with the same accompaniment using lyrics and accompaniment filtering, decodes audio data into mono channels, combines them, and performs energy suppression to create a seamless alternation effect, allowing for automatic synthesis of music with the same accompaniment sung differently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual selection and processing of songs is performed, then synthesis accuracy is improved, but operation complexity and time consumption increase

Engineering Contradiction:
Improvesynthesis accuracyVSAvoidoperation complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically identifies same-source song pairs by extracting and comparing lyrics and accompaniment features without requiring manual selection. The processing module autonomously performs synthesis operations based on automatic identification results, eliminating the need for user intervention in song selection and pairing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes song data by extracting lyrics, identifying accompaniment portions, and storing these features in advance. This preliminary extraction and organization of features enables rapid automatic identification and synthesis execution when needed, improving both accuracy and efficiency.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If manual processing of audio data is performed, then synthesis quality is improved, but resource consumption increases

Engineering Contradiction:
Improvesynthesis qualityVSAvoidresource consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The system extracts and separates lyrics portions from accompaniment portions in songs, storing these extracted features for efficient reuse. By pre-extracting and caching these essential components, the system avoids redundant processing during synthesis operations, reducing computational resource consumption while maintaining synthesis quality.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If automatic synthesis is implemented, then productivity is improved, but system complexity increases

Engineering Contradiction:
Improvesynthesis efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the synthesis process into distinct functional modules: a feature extraction module that preprocesses songs, an identification module that automatically matches same-source pairs, and a synthesis execution module that performs the actual mixing. This modular segmentation enables automatic high-efficiency synthesis while organizing system complexity into manageable, independent components.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3522151B1Method and device for processing dual-source audio data
Publication Date: 2020.11.11 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3522151B1 patent drawingFigure 1~2
  • EP3522151B1 patent drawingFigure 3
  • EP3522151B1 patent drawingFigure 4~5(b)

AI summary

Embodiments of the present invention provide a dual sound source audio data processing method and apparatus. The method includes: obtaining audio data of a same-source song pair, the same-source song pair including two songs having a same accompaniment but sung differently; decoding the audio data of the same-source song pair, to obtain two pieces of mono audio data; combining the two pieces of mono audio data to one piece of two-channel audio data; and dividing play time corresponding to a two-channel audio to multiple play periods, and performing the energy suppression on a left audio channel or a right audio channel of the two-channel audio in different play periods.