Sound Source Separation via Downsampling and Band Extension

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound source separation technologies face challenges in effectively processing mixed sound signals with high-frequency components, leading to increased memory and calculation costs, reduced performance, and difficulty in learning high-resolution sound source separation models, especially in embedded systems and cloud services.

Innovation Solution

A signal processing device and method that applies downsampling processing to mixed sound signals with high-frequency components, generates masks based on the downsampling results, and applies these masks to separate sound sources, utilizing band extension to maintain high-frequency components and reduce processing requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If downsampling processing is applied to mixed sound signals with high-frequency components, then memory and calculation costs are reduced, but high-frequency information may be lost

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidhigh-frequency information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The signal processing is divided into two stages: first, downsampling is applied to reduce the sampling frequency to manageable levels for mask generation; second, band extension is applied to reconstruct high-frequency components. This segmentation allows efficient processing while preserving high-frequency information through the combination of downsampling and subsequent band extension.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Downsampling is performed as a preliminary action before mask generation to reduce computational complexity. The high-frequency information is not permanently lost but is subsequently restored through band extension processing that operates on the downsampled signal and its derived masks.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If high-resolution sound source separation is performed, then sound quality and separation accuracy are maintained, but memory and calculation costs increase significantly

Engineering Contradiction:
Improveseparation accuracyVSAvoidmemory and calculation costs
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The sampling frequency parameter is dynamically adjusted through downsampling to reduce computational complexity during mask generation. This parameter change allows the system to perform separation accuracy calculations at lower computational cost, while band extension subsequently restores the high-frequency content to maintain overall sound quality and separation accuracy.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If downsampling is applied to reduce processing requirements, then memory and calculation costs are reduced, but noise perception may increase in the final output

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidnoise perception
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

Band extension acts as an intermediary process that bridges the downsampled signal and the final high-resolution output. It reconstructs high-frequency components that were removed during downsampling, thereby reducing noise perception in the final output while maintaining the processing efficiency benefits of downsampling throughout the critical mask generation stage.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230419978A1Signal processing device, signal processing method, and program
Publication Date: 2023.12.28 SONY GROUP CORP
  • US20230419978A1 patent drawing
  • US20230419978A1 patent drawing
  • US20230419978A1 patent drawing

AI summary

For example, a signal processing device configured to perform appropriate sound source separation processing is provided. A signal processing device includes: a downconverter configured to apply downsampling processing to a mixed sound signal in which sound source signals included in a high-frequency component higher than a predetermined frequency are mixed; a mask generation unit configured to generate a mask on the basis of a downsampling processing result provided by the downconverter; and a mask processing unit configured to apply the mask generated by the mask generation unit to the mixed sound signal.