Audio event detection method and apparatus, computer device and storage medium

By using an event separation network to generate target masks and feature fusion in audio event detection, the problem of decreased detection performance in complex scenes is solved, and efficient audio event detection is achieved.

CN122455002APending Publication Date: 2026-07-24ZHEJIANG DAHUA TECH CO LTD
0 Cites 0 Cited by

Patent Information

Application Number
CN202610712405.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing audio event detection methods suffer from severe degradation of feature distinguishability in complex scenarios with multi-source interference and low signal-to-noise ratio, leading to a decline in detection performance.

Method used

By acquiring the frequency domain features of the input audio, an event separation network is used to generate a target mask to remove interfering audio. The frequency domain features of the target audio are extracted, and after feature fusion, they are input into the event detection network to output the detection results.

Benefits of technology

It significantly improves the ability to detect audio events in complex scenarios, enhances the signal-to-noise ratio and detection robustness, and is suitable for deployment on edge devices such as mobile phones and IoT devices.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application relates to an audio event detection method, device, computer equipment and storage medium, which comprises the following steps: acquiring frequency domain features of input audio; inputting the frequency domain features into an event separation network to generate a target mask for removing interference audio in the input audio, and obtaining separated target audio frequency domain features according to the target mask and the frequency domain features; performing feature extraction on the frequency domain features and the target audio frequency domain features to obtain fused frequency domain features; inputting the fused frequency domain features into an event detection network to output a detection result representing whether a target audio event exists in the input audio. Through the application, the target audio frequency domain features are first extracted from the input audio, then the frequency domain features and the separated target audio frequency domain features are fused to obtain the fused frequency domain features, and finally the detection result is obtained according to the fused frequency domain features, so that the three-stage series connection structure effectively improves the detection capability of the audio event in a complex scene.
Need to check novelty before this filing date? Find Prior Art