Class-Conditioned Sound Event Localization in Acoustic Mixtures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound event localization and detection (SELD) systems face challenges in accurately localizing sound sources due to movement, obscuration by room reverberation, interference from other sounds, and confusion among sound events, especially when dealing with a large number of classes, and require extensive training data for class-specific models.

Innovation Solution

A class-conditioned SELD system that processes acoustic mixtures using a neural network to output a single ACCDOA vector, identifying the target sound event and determining its direction of arrival (DOA) and distance, even with limited training data, by utilizing a class conditioned SELD network with FiLM blocks and convolution blocks to process spatial and spectral features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a single ACCDOA vector is output for all classes at every time instant, then localization information is provided for all sound events, but the system becomes impractical for large number of classes and computes unnecessary localizations

Engineering Contradiction:
Improvelocalization efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the localization task by introducing class-specific output vectors, where each class has its own ACCDOA representation. This allows the system to process and output localization information selectively for relevant classes rather than computing for all classes uniformly, thereby improving efficiency while managing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by allowing the system to output ACCDOA vectors only for classes that are currently active or relevant at each time instant, rather than computing for all possible classes. This reduces computational load and system complexity while maintaining necessary localization functionality.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If class-specific models are trained to localize only sound events from a single class, then focus is achieved on specific classes, but extremely large training data requirements are needed for each class

Engineering Contradiction:
Improvelocalization accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges multiple class-specific ACCDOA representations into a unified framework that processes all classes simultaneously. By combining class-specific localization vectors with a shared neural network architecture, the system achieves class-specific precision without requiring separate models trained on extremely large datasets for each class.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal SELD system that handles multiple sound event classes through a single integrated model. The class-conditioned ACCDOA representations allow the system to perform class-specific localization tasks universally, reducing the need for extensive class-specific training data while maintaining accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If localization is performed for all classes at every time instant, then complete sound event monitoring is achieved, but computational resources are wasted on classes that are not active or relevant

Engineering Contradiction:
Improvesound event detection completenessVSAvoidcomputational energy
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent introduces dynamic class-conditioned ACCDOA representations that adapt to the current acoustic scene. The system dynamically determines which classes are active or relevant at each time instant and computes localization vectors only for those classes, thereby maintaining complete monitoring reliability while reducing computational energy consumption.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12452590B2Method and system for sound event localization and detection
Publication Date: 2025.10.21 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US12452590B2 patent drawing
  • US12452590B2 patent drawing
  • US12452590B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a system and method for localization of a target sound event. The system collects a first digital representation of an acoustic mixture of sounds of a plurality of sound events, by using an acoustic sensor. The system receives a second digital representation of a sound corresponding to the target sound event. Further, the first digital representation and the second digital representation are processed by a neural network to produce a localization information indicative of a location of an origin of the target sound event with respect to a location of the acoustic sensor.