Audio processing method, apparatus, computing device, storage medium, and program product

The updatable memory bank built through a dynamic memory network solves the problem of insufficient training data for deep learning models in abnormal sound detection, enabling accurate detection of unknown abnormal sounds and autonomous model evolution, thereby improving detection accuracy and adaptability.

CN122290630APending Publication Date: 2026-06-26TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-03-03
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, deep learning models suffer from poor recognition performance in abnormal sound detection due to a lack of training data and the presence of unknown abnormal sounds, making it impossible to effectively identify unexpected abnormal sounds.

Method used

A Dynamic Memory Induction Network (DMIN) is used to construct a dynamically updatable memory bank to store historical audio features and anomaly detection results. Anomaly detection is achieved through feature matching and similarity judgment, and the feature prototypes in the memory bank are dynamically updated during the detection process.

Benefits of technology

It achieves accurate detection of unknown abnormal sounds, reduces reliance on large-scale data, lowers manual and computational costs, and improves the model's autonomous evolution capability and detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290630A_ABST
    Figure CN122290630A_ABST
Patent Text Reader

Abstract

This application provides an audio processing method, apparatus, computing device, storage medium, and program product. By maintaining a memory bank that stores multiple feature prototypes obtained based on audio features from historical audio and anomaly detection results of historical audio, when an audio to be processed is acquired, a target feature prototype matching the audio features of the audio to be processed is determined from the maintained memory bank based on the audio features of the audio to be processed. Then, based on the similarity between the audio features of the audio to be processed and the target feature prototype, the anomaly detection result of the audio to be processed is determined. The feature prototypes stored in the maintained memory bank are updated based on the audio features of the audio to be processed, the similarity between the audio features of the audio to be processed and the target feature prototype, and the anomaly detection result of the audio to be processed. This dynamic updating of the memory bank during inference continuously enhances the recognition capability of the entire system.
Need to check novelty before this filing date? Find Prior Art