Audio event detection method and apparatus, computer device and storage medium
By using an event separation network to generate target masks and feature fusion in audio event detection, the problem of decreased detection performance in complex scenes is solved, and efficient audio event detection is achieved.
Patent Information
- Application Number
- CN202610712405.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-07-24
AI Technical Summary
Existing audio event detection methods suffer from severe degradation of feature distinguishability in complex scenarios with multi-source interference and low signal-to-noise ratio, leading to a decline in detection performance.
By acquiring the frequency domain features of the input audio, an event separation network is used to generate a target mask to remove interfering audio. The frequency domain features of the target audio are extracted, and after feature fusion, they are input into the event detection network to output the detection results.
It significantly improves the ability to detect audio events in complex scenarios, enhances the signal-to-noise ratio and detection robustness, and is suitable for deployment on edge devices such as mobile phones and IoT devices.