Event Mask Generation for Audio Signal Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies struggle to detect nonstationary sounds with unknown spectrum shapes as audio events, leading to incomplete detection and omission of audio events.
Innovation Solution
A mask generation device and method that extracts sound pressure information from a spectrogram and applies a binarization process to generate an event mask, indicating the time period of an audio event, allowing for the detection of audio events with unknown spectrum shapes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing voice activity detection methods are used to distinguish voice from noise, then voice sections can be identified, but nonstationary sounds with unknown spectrum shapes cannot be detected as audio events
Solution Approach 1:
The invention changes the detection parameter from assuming a specific spectrum shape (voice) to detecting sound pressure level changes regardless of spectrum shape. By using binarization of sound pressure information extracted from spectrogram, the system can detect both stationary and nonstationary sounds with any spectrum shape as audio events, resolving the contradiction between adaptability to unknown spectra and reliability in detection accuracy.
2Measurement precision
If temporal waveform is used to determine sound pressure, then voice detection is possible, but sounds with strong power in specific frequencies but unknown spectrum shape cannot be sufficiently detected
Solution Approach 1:
The invention introduces spectrogram as an intermediary representation between the temporal waveform and the sound detection. The spectrogram provides frequency-domain information that allows extraction of sound pressure information for each frequency component. By binarizing this extracted information, the system achieves both precise sound pressure measurement and adaptability to detect frequency-specific nonstationary sounds with unknown spectrum shapes.
3Reliability
If event mask is applied to spectrogram to reduce noise, then voice recognition accuracy improves, but nonstationary audio events with unknown spectrum shapes are omitted
Solution Approach 1:
Instead of applying a fixed event mask to suppress noise and potentially losing information, the invention inverts the approach by first extracting sound pressure information from the spectrogram and then applying binarization to create an adaptive event mask. This inverted approach ensures that both stationary and nonstationary sounds are detected as audio events, preventing information loss while maintaining noise reduction capabilities for voice recognition.
Data Source
AI summary
An extraction unit extracts sound pressure information from a spectrogram. A binarization unit carries out binarization on the extracted sound pressure information in order to generate an event mask indicating a time period in which an audio event exists.


