Sound Source Separation via Dynamic Amplitude Spectrum Combination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound source separation technologies, such as Multi Channel Wiener Filter (MWF)-based methods using Deep Neural Networks (DNNs), face challenges in achieving high separation performance due to limited learning data and complexity, leading to errors in amplitude spectrum estimation and subsequent deterioration in separation performance.

Innovation Solution

The approach combines different sound source separation systems with varying time characteristics, utilizing amplitude spectrum estimation algorithms like DNN and LSTM, which differ in output characteristics, to enhance separation performance by linearly combining their estimation results and dynamically adjusting combination parameters based on evaluation functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If DNN-based amplitude spectrum estimation is used to improve sound source separation performance, then separation accuracy increases, but errors in amplitude spectrum estimation occur due to limited learning data and problem complexity

Engineering Contradiction:
Improveamplitude spectrum estimation accuracyVSAvoidseparation performance stability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines multiple sound source separation systems with different time characteristics (DNN, LSTM, and other algorithms) into a unified framework. By merging the estimation results of multiple systems through linear combination with dynamically adjusted parameters, the invention achieves more reliable and stable separation performance while maintaining high accuracy, effectively resolving the contradiction between estimation accuracy and performance stability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces dynamic parameter adjustment mechanisms where combination parameters are adapted based on evaluation functions that assess the performance of individual systems in different time frames. This dynamic adaptation allows the system to optimize the contribution of each algorithm according to current conditions, improving both accuracy and reliability simultaneously.

Inventive Principle:
Principle #15Dynamics

2Reliability

If multiple sound source separation systems are combined to improve separation performance, then noise and artifacts are reduced, but system complexity increases

Engineering Contradiction:
Improveseparation performance stabilityVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple sound source separation systems into a single integrated framework that processes audio signals through parallel algorithmic paths. By combining DNN, LSTM, and other separation systems with a unified linear combination mechanism and dynamic parameter adjustment, the invention achieves enhanced reliability while managing complexity through systematic integration rather than separate independent systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal sound source separation framework that can accommodate multiple different algorithms (DNN, LSTM, and others) through a common architecture. This multi-functional system uses a standardized linear combination mechanism and evaluation function approach that works across different algorithm types, reducing overall system complexity while maintaining the benefits of multiple specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If DNN learning is performed with limited learning data to reduce error, then amplitude spectrum accuracy improves, but learning difficulty increases due to problem complexity

Engineering Contradiction:
Improveamplitude spectrum estimation accuracyVSAvoidlearning process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple learning systems (DNN, LSTM, and other algorithms) that have been trained on limited learning data. By merging their estimation results through linear combination with dynamically adjusted parameters, the invention achieves high amplitude spectrum estimation accuracy while distributing the learning complexity across multiple specialized systems rather than requiring one extremely complex system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs multiple algorithms with different strengths and time characteristics, where each algorithm performs partial learning on the limited data. By combining these partial results, the system achieves comprehensive accuracy that would be difficult for any single algorithm to achieve alone, effectively managing learning complexity through distributed specialization.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3511937B1Device and method for sound source separation, and program
Publication Date: 2023.08.23 SONY GROUP CORP
  • EP3511937B1 patent drawingFigure 1
  • EP3511937B1 patent drawingFigure 2
  • EP3511937B1 patent drawingFigure 3

AI summary

The present technology relates to a sound source separation device that enables to achieve higher separation performance, and a method and a program. The sound source separation device includes a combining unit that combines a first sound source separation signal of a predetermined sound source, the first sound source separation signal being separated from a mixed sound signal by a first sound source separation system, with a second sound source separation signal of the sound source, the second sound source separation signal being separated from the mixed sound signal by a second sound source separation system that differs in separation performance from the first sound source separation system in predetermined units of time, and that outputs a sound source separation signal obtained by the combination. The present technology can be applied to the sound source separation device.