Spatial-Region Neural Feature Extraction for Flexible Sound Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network-based audio processing systems struggle with efficiently separating multiple localized sound sources due to the need for predefined spatial angles and computational inefficiencies, lacking versatility and flexibility in source selection.
Innovation Solution
A neural network architecture that incorporates spatial region-dependent feature extraction, allowing user-defined target regions at runtime, reduces computational redundancy by focusing on a specified spatial range, and integrates DOA-dependent layers to enhance separation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the neural network is trained to extract signal components from all directions simultaneously, then the system can handle any target direction, but computational resources are wasted due to redundancy in output streams
Solution Approach 1:
The system dynamically adapts the neural network's spatial processing based on the target direction parameter provided at runtime. Instead of a fixed architecture that processes all directions, the system configures the spatial processing dynamically to focus only on the relevant direction, reducing computational redundancy while maintaining versatility.
Solution Approach 2:
The patent applies local quality by making the neural network's spatial processing specialized for a specific target direction rather than uniformly processing all directions. The spatial processing is localized to the region of interest, eliminating redundant computations for irrelevant directions while preserving the ability to handle any direction by changing the target parameter.
2Productivity
If the neural network is trained for a fixed spatial angle, then computational resources are used efficiently, but the system lacks flexibility in selecting different angles at runtime
Solution Approach 1:
The system achieves dynamic adaptability by accepting a target direction as a runtime parameter and configuring the spatial processing accordingly. This allows the system to maintain computational efficiency for a specific angle while being flexible in selecting different angles when needed, resolving the contradiction between efficiency and adaptability.
3Adaptability or versatility
If multiple output streams are generated for all directions, then complete coverage of all possible targets is achieved, but the majority of output streams are unused and create redundancy
Solution Approach 1:
The patent extracts only the relevant spatial information corresponding to the target direction from the full set of possible directions. By taking out only the necessary component (target direction) and discarding the rest (other directions), the system eliminates redundancy while maintaining the ability to handle any target by changing the extraction parameter.
Data Source
Figure 1~2
Figure 3~4A
Figure 4B
AI summary
A method for estimating target sound sources located in at least one target spatial region among a plurality of spatial regions by receiving a plurality of signals associated with one of a plurality of microphone signals comprising sound events generated by the plurality of sound sources, extracting via a neural network a plurality of features obtained by training the neural network for a different spatial region among the plurality of spatial regions, generating, by the processor, another plurality of features based on the extracted plurality of features wherein the another plurality of features corresponds to the at least one target spatial region, detecting or estimating, by the processor, at least one sound source among the target sound sources in the target spatial region based on the another plurality of features corresponding to the at least one target spatial region.