一种基于时频自适应与协同注意力的声音事件分类方法

By employing time-frequency adaptive and collaborative attention mechanisms, a sound event classification model is constructed, which solves the problem of low classification accuracy of existing models in complex acoustic scenarios. This model improves classification accuracy and modeling ability without increasing the number of parameters, and enhances the feature extraction capability for complex acoustic scenarios.

CN122417082APending Publication Date: 2026-07-17JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610850070.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing sound event classification models are not accurate enough in complex acoustic scenarios, and it is difficult to improve global and local modeling capabilities without increasing the number of parameters. Furthermore, existing attention mechanisms are difficult to balance computational efficiency with long-distance dependency modeling capabilities.

Method used

Employing a time-frequency adaptive and collaborative attention mechanism, a sound event classification model is constructed using a time-frequency decoupled adaptive module (TFDA), a cross-scale collaborative attention module, a gating module, a multilayer perceptron (MLP), and a global average pooling (GAP). Combined with time-frequency separation dynamic convolution and a multi-scale expanded deep residual network, efficient feature extraction of complex acoustic patterns is achieved.

Benefits of technology

Without significantly increasing the number of parameters, it improves the accuracy of sound event classification and global and local modeling capabilities, and enhances the feature representation capabilities of complex acoustic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417082A_ABST
    Figure CN122417082A_ABST
Patent Text Reader

Abstract

本申请提供的一种基于时频自适应与协同注意力的声音事件分类方法,其构建了时频分离动态卷积模块TFSD‑Conv通过时域动态卷积与频域动态卷积对声学特征进行解耦建模,并结合多尺度扩张深度残差网络MSD‑Resnet扩大网络感受野,从而实现对复杂声学模式的高效自适应的特征提取。
Need to check novelty before this filing date? Find Prior Art