Energy-Based Label Encoding for Polyphonic Sound Event Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing label encoding methods for sound event recognition models fail to accurately capture multi-sound event components, leading to errors and limited recognition performance in deep neural network-based sound event recognition systems.
Innovation Solution
A method and apparatus for label encoding based on energy information, which identifies event intervals in sound signals, separates sound sources, determines energy information for each sound event, and performs label encoding using a sum, scale factor, and bias to improve learning and recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If one-hot label encoding is used for sound events, then the encoding process is simple, but recognition performance is limited due to inability to capture multi-sound event components
Solution Approach 1:
The patent segments multi-sound event intervals into individual sound event components using sound source separation. Each sound event is processed separately to extract its energy information, then recombined to create comprehensive ground truth labels that preserve both individual and mixed sound event characteristics.
Solution Approach 2:
The patent introduces energy information as a new parameter for label encoding. Instead of binary one-hot encoding, it uses energy values extracted from sound source separation to create continuous label representations that capture the intensity and characteristics of each sound event component.
2Measurement precision
If sound source separation is performed on multi-sound event intervals, then energy information for each sound event can be determined, but the processing complexity increases
Solution Approach 1:
The patent performs sound source separation and energy information extraction in advance during the ground truth preparation phase. This preliminary processing creates pre-computed energy labels that can be directly used for training without requiring complex real-time processing during model inference.
Solution Approach 2:
The patent introduces energy information as an intermediary representation between raw audio signals and final labels. This intermediary captures the essential characteristics of each sound event component, simplifying the subsequent label encoding process while maintaining accuracy.
3Ease of manufacture
If ground truth information is obtained using traditional one-hot encoding, then the label encoding process is straightforward, but errors occur in multi-sound event intervals
Solution Approach 1:
The patent uses sound source separation results as feedback to refine the label encoding process. By analyzing the separated sound event components and their energy distributions, the system generates corrected ground truth labels that accurately represent both individual and mixed sound events in overlapping intervals.
Solution Approach 2:
The patent creates composite label representations by combining energy information from multiple separated sound event components. This composite approach preserves the characteristics of individual sound events while also capturing their mixed occurrence, creating more reliable ground truth for training.
Data Source
AI summary
Disclosed is a method and apparatus for label encoding in a multi-sound event interval. The method includes identifying an event interval in which a plurality of sound events occurs in a sound signal, separating a sound source into sound event signals corresponding to each sound event by performing sound source separation on the event interval, determining energy information for each of the sound event signals, and performing label encoding based on the energy information.


