Sound Source Separation via Downsampling and Band Extension
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sound source separation technologies face challenges in effectively processing mixed sound signals with high-frequency components, leading to increased memory and calculation costs, reduced performance, and difficulty in learning high-resolution sound source separation models, especially in embedded systems and cloud services.
Innovation Solution
A signal processing device and method that applies downsampling processing to mixed sound signals with high-frequency components, generates masks based on the downsampling results, and applies these masks to separate sound sources, utilizing band extension to maintain high-frequency components and reduce processing requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If downsampling processing is applied to mixed sound signals with high-frequency components, then memory and calculation costs are reduced, but high-frequency information may be lost
Solution Approach 1:
The signal processing is divided into two stages: first, downsampling is applied to reduce the sampling frequency to manageable levels for mask generation; second, band extension is applied to reconstruct high-frequency components. This segmentation allows efficient processing while preserving high-frequency information through the combination of downsampling and subsequent band extension.
Solution Approach 2:
Downsampling is performed as a preliminary action before mask generation to reduce computational complexity. The high-frequency information is not permanently lost but is subsequently restored through band extension processing that operates on the downsampled signal and its derived masks.
2Measurement precision
If high-resolution sound source separation is performed, then sound quality and separation accuracy are maintained, but memory and calculation costs increase significantly
Solution Approach 1:
The sampling frequency parameter is dynamically adjusted through downsampling to reduce computational complexity during mask generation. This parameter change allows the system to perform separation accuracy calculations at lower computational cost, while band extension subsequently restores the high-frequency content to maintain overall sound quality and separation accuracy.
3Productivity
If downsampling is applied to reduce processing requirements, then memory and calculation costs are reduced, but noise perception may increase in the final output
Solution Approach 1:
Band extension acts as an intermediary process that bridges the downsampled signal and the final high-resolution output. It reconstructs high-frequency components that were removed during downsampling, thereby reducing noise perception in the final output while maintaining the processing efficiency benefits of downsampling throughout the critical mask generation stage.
Data Source
AI summary
For example, a signal processing device configured to perform appropriate sound source separation processing is provided. A signal processing device includes: a downconverter configured to apply downsampling processing to a mixed sound signal in which sound source signals included in a high-frequency component higher than a predetermined frequency are mixed; a mask generation unit configured to generate a mask on the basis of a downsampling processing result provided by the downconverter; and a mask processing unit configured to apply the mask generated by the mask generation unit to the mixed sound signal.


