AI Sound Source Separation via Dictionary Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source separation technologies are limited in the number of sound sources they can separate based on the number of microphones used and struggle with noise removal and voice separation in real environments, particularly when multiple sound sources overlap.
Innovation Solution
A method and device using dictionary learning with the K-SVD algorithm for sound source separation, transforming sound data into mel-scale, and non-negative matrix factorization to separate and detect target sound sources from overlapping sound sources, with a gain matrix update to minimize differences and determine successful detection based on threshold values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional sound source separation methods using multiple microphones are used, then the number of separable sound sources is limited by the number of microphones, but the device complexity increases
Solution Approach 1:
The patent replaces the mechanical approach of using multiple microphones with an artificial intelligence-based signal processing system. The sound source separation is achieved through AI algorithms that analyze audio signals and separate sound sources computationally, rather than relying on physical microphone arrays. This substitution reduces hardware complexity while maintaining or improving separation capability.
Solution Approach 2:
The patent transforms sound data into mel-scale representation and uses dictionary learning with K-SVD algorithm to change the parameter space for analysis. By transforming the audio data into different feature domains (time-frequency representation, mel-scale), the system can separate sound sources more effectively without requiring additional physical sensors.
2Reliability
If traditional sound source separation methods are used, then previously learned sound sources can be separated, but the adaptability to various real environment sound sources deteriorates
Solution Approach 1:
The patent performs dictionary learning in advance to build a comprehensive database of sound source characteristics. By pre-learning various sound source patterns and creating dictionaries for different types of sounds, the system prepares beforehand to handle diverse real-world scenarios. This preliminary action enables the system to adapt to various sound sources in real environments without requiring retraining for each specific case.
Solution Approach 2:
The patent creates a universal sound source separation system that can handle multiple types of sound sources through a single AI model. The dictionary learning approach builds generalized representations that work across different sound sources, environments, and conditions, making the system versatile rather than specialized for single sound sources.
3Reliability
If sound source separation is performed in overlapping sound sources, then noise removal and voice separation improve, but the measurement precision deteriorates due to signal mixing
Solution Approach 1:
The patent segments the overlapping sound source signal into multiple components through AI-based separation. The system divides the mixed audio signal into individual sound source streams by analyzing temporal, spectral, and statistical characteristics. This segmentation allows the system to extract and process each sound source separately, maintaining measurement precision even when sources overlap in time and frequency domains.
Data Source
AI summary
An artificial sound source separation method and device are disclosed. The sound source separation method by the artificial sound source separation device based on dictionary learning generates a dictionary matrix by performing dictionary learning, receives an overlapping sound source in which at least two sound sources are mixed and separates a target sound source from the overlapping sound source based on the dictionary matrix; and detecting the target sound source. The dictionary learning may be performed using a K-SVD algorithm. The intelligent computing device configuring a sound source processing device of the present disclosure may be associated with an artificial intelligence module, drone (unmanned aerial vehicle, UAV), robot, augmented reality (AR) devices, virtual reality (VR) devices, devices related to 5G services, and the like.


