Class-Conditioned Sound Event Localization in Acoustic Mixtures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound event localization and detection (SELD) systems face challenges in accurately localizing sound sources due to movement, obscuration by room reverberation, interference from other sounds, and confusion among sound events, especially when dealing with a large number of classes, and require extensive training data for class-specific models.
Innovation Solution
A class-conditioned SELD system that processes acoustic mixtures using a neural network to output a single ACCDOA vector, identifying the target sound event and determining its direction of arrival (DOA) and distance, even with limited training data, by utilizing a class conditioned SELD network with FiLM blocks and convolution blocks to process spatial and spectral features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a single ACCDOA vector is output for all classes at every time instant, then localization information is provided for all sound events, but the system becomes impractical for large number of classes and computes unnecessary localizations
Solution Approach 1:
The patent segments the localization task by introducing class-specific output vectors, where each class has its own ACCDOA representation. This allows the system to process and output localization information selectively for relevant classes rather than computing for all classes uniformly, thereby improving efficiency while managing complexity.
Solution Approach 2:
The patent implements partial action by allowing the system to output ACCDOA vectors only for classes that are currently active or relevant at each time instant, rather than computing for all possible classes. This reduces computational load and system complexity while maintaining necessary localization functionality.
2Measurement precision
If class-specific models are trained to localize only sound events from a single class, then focus is achieved on specific classes, but extremely large training data requirements are needed for each class
Solution Approach 1:
The patent merges multiple class-specific ACCDOA representations into a unified framework that processes all classes simultaneously. By combining class-specific localization vectors with a shared neural network architecture, the system achieves class-specific precision without requiring separate models trained on extremely large datasets for each class.
Solution Approach 2:
The patent creates a universal SELD system that handles multiple sound event classes through a single integrated model. The class-conditioned ACCDOA representations allow the system to perform class-specific localization tasks universally, reducing the need for extensive class-specific training data while maintaining accuracy.
3Reliability
If localization is performed for all classes at every time instant, then complete sound event monitoring is achieved, but computational resources are wasted on classes that are not active or relevant
Solution Approach 1:
The patent introduces dynamic class-conditioned ACCDOA representations that adapt to the current acoustic scene. The system dynamically determines which classes are active or relevant at each time instant and computes localization vectors only for those classes, thereby maintaining complete monitoring reliability while reducing computational energy consumption.
Data Source
AI summary
Embodiments of the present disclosure disclose a system and method for localization of a target sound event. The system collects a first digital representation of an acoustic mixture of sounds of a plurality of sound events, by using an acoustic sensor. The system receives a second digital representation of a sound corresponding to the target sound event. Further, the first digital representation and the second digital representation are processed by a neural network to produce a localization information indicative of a location of an origin of the target sound event with respect to a location of the acoustic sensor.


