Sound event detection method and system with class distribution and temporal context collaborative cues

By using a method of global distribution and local temporal collaborative prompts, the audio pre-trained model is fine-tuned, which solves the problem of insufficient modeling of global features and local temporal features in sound event detection. This achieves a high-efficiency improvement in audio classification and localization performance, and is suitable for practical scenarios such as intelligent monitoring and environmental perception.

CN120048284BActive Publication Date: 2026-03-03JIANGSU UNIV
2 Cites 0 Cited by

Patent Information

Application Number
CN202510211179.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-03-03
Estimated Expiration
2045-02-25

Smart Images

  • Figure CN120048284B_ABST
    Figure CN120048284B_ABST
Patent Text Reader

Abstract

The application discloses a sound event detection method and system based on distribution and timing context cooperative prompting, converts an original audio signal into a signal frame sequence, extracts a mel filter bank feature through a pre-training model branch, and extracts a mel spectrum feature through a downstream model branch; a combination of a local timing prompt module and a global distribution prompt module is introduced into each layer of the pre-training model, the pre-training model outputs an audio sequence feature and the global distribution prompt module; the audio sequence and the output of the downstream model are fused in features, frame-level prediction probabilities of the downstream model are calculated, and an audio positioning task is realized; then, the frame-level prediction probabilities of the downstream model are used to generate sentence-level prediction probabilities; for the pre-training model, the global distribution prompt module is processed to obtain sentence-level prediction probabilities of the pre-training model; finally, the two sentence-level prediction probabilities are fused to obtain a classification result of a sound event. The application can significantly improve the audio classification and positioning performance of sound event detection.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Weak supervision sound event detection method and system based on adaptive hierarchical aggregation

    CN114974303A

  • Audio detection model training method, audio detection method and related device

    CN117059075A