Scene-Adaptive Sound Source Separation Using Embedding Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound source separation technologies struggle to effectively isolate a target sound source from a mixed sound source, particularly in environments with background noise, using conventional methods that lack scene-specific adaptation.

Innovation Solution

A method involving scene information-based sound source separation, utilizing a pre-training process to convert a first embedding vector into a second embedding vector, allowing for the separation of a target sound source from a mixed sound source by leveraging scene-specific information and neural network processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional sound source separation methods are used, then the separation process is simple, but the accuracy of isolating target sound sources in noisy environments is poor

Engineering Contradiction:
Improveaccuracy of sound source separationVSAvoidcomplexity of separation process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs pre-training processes before actual sound source separation to learn scene information and convert embedding vectors. This preliminary learning phase enables the model to adapt to different acoustic environments in advance, improving separation accuracy when processing actual mixed sound sources.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system converts a first embedding vector into a second embedding vector based on scene information, effectively changing the parameter representation. This vector transformation adapts the sound source characteristics to specific acoustic scenes, enabling more accurate separation in different environments.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If scene-specific adaptation is implemented, then the accuracy of sound source separation improves, but the processing time and computational complexity increase

Engineering Contradiction:
Improveaccuracy of sound source separationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Scene information learning and embedding vector conversion are performed in advance through pre-training processes. This allows the system to quickly adapt to new scenes without extensive processing time during actual sound source separation operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses embedding vectors as compressed representations of sound source characteristics. Instead of processing raw audio data extensively, the model works with these compact vector representations, significantly reducing computational time while maintaining separation accuracy.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple pre-training processes are performed, then the adaptability to different acoustic scenes improves, but the computational resources and processing steps increase

Engineering Contradiction:
Improveadaptability to acoustic scenesVSAvoidnumber of processing steps
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The pre-training process is divided into distinct stages: first pre-training for scene information learning, and second pre-training for embedding vector conversion. This segmentation allows the system to tackle different aspects of adaptability in separate, manageable steps rather than attempting to learn everything simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Embedding vectors serve as an intermediary representation between raw audio data and the final sound source separation output. These vectors capture essential scene-specific characteristics in a compressed form, enabling the model to adapt to different acoustic environments without requiring extensive processing of raw audio data for each scene.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12424241B2Method for separating target sound source from mixed sound source and electronic device thereof
Publication Date: 2025.09.23 SAMSUNG ELECTRONICS CO LTD
  • US12424241B2 patent drawing
  • US12424241B2 patent drawing
  • US12424241B2 patent drawing

AI summary

A method for separating a target sound source, includes: obtaining a mixed sound source including at least one sound source; obtaining, based on the mixed sound source, scene information related to the mixed sound source; converting, based on the scene information, a first embedding vector corresponding to a designated sound source group into a second embedding vector; and separating, based on the mixed sound source and the second embedding vector, the target sound source from the mixed sound source.