Sound Source Separation via Dynamic Amplitude Spectrum Combination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source separation technologies, such as Multi Channel Wiener Filter (MWF)-based methods using Deep Neural Networks (DNNs), face challenges in achieving high separation performance due to limited learning data and complexity, leading to errors in amplitude spectrum estimation and subsequent deterioration in separation performance.
Innovation Solution
The approach combines different sound source separation systems with varying time characteristics, utilizing amplitude spectrum estimation algorithms like DNN and LSTM, which differ in output characteristics, to enhance separation performance by linearly combining their estimation results and dynamically adjusting combination parameters based on evaluation functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If DNN-based amplitude spectrum estimation is used to improve sound source separation performance, then separation accuracy increases, but errors in amplitude spectrum estimation occur due to limited learning data and problem complexity
Solution Approach 1:
The patent combines multiple sound source separation systems with different time characteristics (DNN, LSTM, and other algorithms) into a unified framework. By merging the estimation results of multiple systems through linear combination with dynamically adjusted parameters, the invention achieves more reliable and stable separation performance while maintaining high accuracy, effectively resolving the contradiction between estimation accuracy and performance stability.
Solution Approach 2:
The patent introduces dynamic parameter adjustment mechanisms where combination parameters are adapted based on evaluation functions that assess the performance of individual systems in different time frames. This dynamic adaptation allows the system to optimize the contribution of each algorithm according to current conditions, improving both accuracy and reliability simultaneously.
2Reliability
If multiple sound source separation systems are combined to improve separation performance, then noise and artifacts are reduced, but system complexity increases
Solution Approach 1:
The patent merges multiple sound source separation systems into a single integrated framework that processes audio signals through parallel algorithmic paths. By combining DNN, LSTM, and other separation systems with a unified linear combination mechanism and dynamic parameter adjustment, the invention achieves enhanced reliability while managing complexity through systematic integration rather than separate independent systems.
Solution Approach 2:
The patent creates a universal sound source separation framework that can accommodate multiple different algorithms (DNN, LSTM, and others) through a common architecture. This multi-functional system uses a standardized linear combination mechanism and evaluation function approach that works across different algorithm types, reducing overall system complexity while maintaining the benefits of multiple specialized systems.
3Measurement precision
If DNN learning is performed with limited learning data to reduce error, then amplitude spectrum accuracy improves, but learning difficulty increases due to problem complexity
Solution Approach 1:
The patent combines multiple learning systems (DNN, LSTM, and other algorithms) that have been trained on limited learning data. By merging their estimation results through linear combination with dynamically adjusted parameters, the invention achieves high amplitude spectrum estimation accuracy while distributing the learning complexity across multiple specialized systems rather than requiring one extremely complex system.
Solution Approach 2:
The patent employs multiple algorithms with different strengths and time characteristics, where each algorithm performs partial learning on the limited data. By combining these partial results, the system achieves comprehensive accuracy that would be difficult for any single algorithm to achieve alone, effectively managing learning complexity through distributed specialization.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present technology relates to a sound source separation device that enables to achieve higher separation performance, and a method and a program. The sound source separation device includes a combining unit that combines a first sound source separation signal of a predetermined sound source, the first sound source separation signal being separated from a mixed sound signal by a first sound source separation system, with a second sound source separation signal of the sound source, the second sound source separation signal being separated from the mixed sound signal by a second sound source separation system that differs in separation performance from the first sound source separation system in predetermined units of time, and that outputs a sound source separation signal obtained by the combination. The present technology can be applied to the sound source separation device.