Text-Guided Sound Extraction from Noisy Mixture Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound extraction methods struggle to accurately extract specific sounds from mixed sound sources due to environmental noise interference, and they fail when the desired sound range does not match predefined event types, leading to incomplete extraction.
Innovation Solution
A sound extraction system that uses a learning subsystem to generate models for extracting target sounds based on mixture signals and variable-length onomatopoeia texts, employing feature extraction, text-embedded extraction, and time-frequency mask generation to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If predefined event types are used for sound classification, then sound extraction can be performed using fixed categories, but sounds with granularity finer than predefined types cannot be accurately extracted
Solution Approach 1:
The patent transforms the static predefined event type classification into a dynamic text-based range specification system. Users can flexibly define sound extraction ranges using natural language text descriptions rather than being constrained by fixed event categories. This allows the system to adapt to various granularity levels and specific sound ranges that users need, resolving the contradiction between ease of operation and extraction accuracy.
Solution Approach 2:
The patent changes the parameter for defining sound ranges from discrete predefined event types to continuous text-based specifications. By using text to represent sound ranges, the system can capture subtle variations in sound characteristics and time-frequency domains, enabling precise extraction of sounds with finer granularity while maintaining ease of use through natural language input.
2Device complexity
If conventional sound extraction methods are used, then processing can be performed with simple techniques, but extraction accuracy is significantly reduced by environmental noise
Solution Approach 1:
The patent introduces text-based range specifications as an intermediary between the user's extraction intent and the actual sound processing. This text representation serves as a flexible mediator that can precisely define sound ranges in both time and frequency domains, allowing the system to accurately isolate target sounds from environmental noise while maintaining reasonable processing complexity through learned models.
3Stability of the object's composition
If fixed event type classification is applied, then sound categories are clearly defined, but the system cannot adapt to user-specific extraction needs
Solution Approach 1:
The patent creates a universal text-based interface that can serve multiple sound extraction needs across different applications and users. Rather than requiring separate predefined event type systems for different scenarios, the text-based range specification provides a single versatile mechanism that adapts to various user requirements, maintaining stability through consistent processing while achieving versatility through flexible text input.
Data Source
AI summary
To provide a sound extraction system and a sound extraction method capable of accurately extracting, from mixture signals, a signal corresponding to a sound which a user wants to extract. The sound extraction system includes a sound extraction device configured to extract, from mixture signals including a signal corresponding to an extraction target sound, the signal corresponding to the extraction target sound. The sound extraction device is configured to extract the signal corresponding to the extraction target sound from the mixture signals based on the mixture signals and a text representing a range of the extraction target sound.


