Sound Source Separation via Hidden Variable Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for separating mixed sound signals, particularly in pop music, into human voice and accompaniment are inefficient, as they rely on end-to-end deterministic models that struggle to accurately isolate sound sources, impacting music editing and retrieval processes.

Innovation Solution

A method and apparatus that utilize a coding model to extract hidden variables representing human voice and accompaniment features from mixed sound signals, followed by decoding models to separate these components, employing neural networks trained with loss functions to minimize errors and improve sound source separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If an end-to-end deterministic model is used to separate voice and accompaniment, then the separation process is simplified, but the accuracy of sound source isolation deteriorates

Engineering Contradiction:
Improveseparation process complexityVSAvoidsound source isolation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the separation process into two distinct stages: first extracting hidden variables from the mixed sound signal using a coding model, then separating the actual sound sources using decoding models. This segmentation allows each stage to focus on specific tasks, improving overall separation accuracy while maintaining manageable complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces hidden variables as an intermediary representation between the mixed sound signal and the final separated sources. These hidden variables capture essential features of the sound sources without directly revealing them, enabling more accurate separation through the decoding models while keeping the process structured

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If model training is performed online, then real-time adaptation is improved, but resource consumption and processing time increase

Engineering Contradiction:
Improvereal-time adaptation capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent performs model training in advance (offline) before actual separation tasks. The coding and decoding models are pre-trained on training datasets to learn the characteristics of different sound sources. This preliminary action allows the models to be ready for deployment without requiring continuous online training, reducing real-time computational burden while maintaining adaptability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11430427B2Method and electronic device for separating mixed sound signal
Publication Date: 2022.08.30 BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
  • US11430427B2 patent drawing
  • US11430427B2 patent drawing
  • US11430427B2 patent drawing

AI summary

This application can provide a method and electronic device for separating mixed sound signal. The method includes: obtaining a first hidden variable representing a human voice feature and a second hidden variable representing an accompaniment sound feature by inputting feature data of a mixed sound extracted from a mixed sound signal into a coding model for the mixed sound; obtaining first feature data of a human voice and second feature data of an accompaniment sound by inputting the first hidden variable and the second hidden variable into a first decoding model for the human voice and a second decoding model for the accompaniment sound respectively; and obtaining, based on the first feature data and the second feature data, the human voice and the accompaniment sound.