Indoor abnormal sound recognition method based on L-Gtg features and improved FastViT model

Through the L-Gtg feature and the improved DFHL-FastViT model, the accuracy of abnormal sound recognition in indoor high noise environments is solved. The dual-channel feature fusion and HiLo attention module are adopted to improve the recognition accuracy and noise immunity and reduce the computational complexity.

CN120260620APending Publication Date: 2025-07-04GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271584.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art abnormal sound recognition in indoor environments is affected by high-level background noise, has poor performance and poor noise resistance.

Method used

The L-Gtg feature and the improved DFHL-FastViT model are adopted, and the dual-channel feature fusion module and the HiLo attention module are used to enhance the model's capture ability of time-frequency features and reduce the computational complexity.

Benefits of technology

High recognition accuracy and robustness are achieved in high noise environments, with a reduced parameter volume of 5.39%, and improved inference speed, suitable for real-time processing of abnormal sound events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260620A_ABST
    Figure CN120260620A_ABST
Patent Text Reader

Abstract

The invention aims to provide an indoor abnormal sound recognition method based on L-Gtg features and an improved FastViT model in order to solve the problem that abnormal sound recognition is affected by a masking effect generated under a high-level background noise condition in an indoor environment. The method comprises the following steps: firstly, extracting a Log-Mel feature and a Gammatone feature from an original audio to form a new L-Gtg feature, and increasing the diversified representation of the features; then, a DFHL-FastViT audio classification model is proposed, a dual-channel feature fusion module is designed, weights are dynamically allocated to different feature channels of L-Gtg to be fused, a HiLo attention module is introduced to improve an MHSA attention module in FastViT, the global relation and local sudden change in time-frequency features are captured, and meanwhile the calculation cost and the overall complexity of the model are reduced. Experimental results show that the method provided by the invention has certain noise resistance while keeping robust identification capability with high accuracy in a high indoor background noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal sound detection, and particularly to an indoor abnormal sound recognition method based on L-Gtg features and an improved FastViT model. Background Art

[0002] In recent years, with the rapid development of Internet of Things technology, abnormal sound event detection has become an important research direction in the field of audio signal processing and plays an important role in practical applications such as industrial monitoring, security monitoring, fault diagnosis, and emergency response. Compared with video-based abnormal event detection methods, sound detection has significant advantages, including lower computational resource requirements, smaller storage requirements, and the ability to work efficiently in low-light or occluded environments. However, the indoor environment is usually accompanied by complex background noises, such as the noise generated by human activities, continuous machine operation sounds, etc., which pose severe challenges to the abnormal sound detection task. However, with the development of deep learning technology, new solutions have been provided for abnormal sound detection.

[0003] Currently, in existing research, Log-Mel spectrum is used as the time-frequency feature representation of audio signals, and convolutional neural networks are combined to improve the accuracy of environmental sound recognition. However, due to the influence of indoor noise and the poor anti-noise performance of this scheme, the performance of its abnormal sound recognition is not good. Summary of the Invention

[0004] To solve the deficiencies of the prior art and meet the anti-noise requirements of abnormal sound recognition in indoor environments, the present invention proposes a detection method based on L-gtg features and an improved FastViT model, which effectively weakens the masking effect of indoor background noise, highlights the features of instantaneously generated abnormal sounds, and realizes more accurate and rapid recognition of abnormal sounds in complex indoor noise environments.

[0005] To achieve this purpose, the present invention provides the following technical solutions:

[0006] Step S1: Select 10 types of abnormal sound categories that meet indoor scenarios from the audio samples in the public datasets (ESC50, Acoustic Event Dataset (AED)) to form an indoor abnormal sound dataset. And select restaurant noise, fan sound, and kitchen frying sound from the Noise-x92 dataset and the Disco dataset

[13] as background noises, and mix them with abnormal sounds under different signal-to-noise ratio (SNR) conditions (-20db, -15db, -10db, 0db) to generate a noisy indoor abnormal sound dataset. The total number of samples in the dataset is 5538 audio samples. Then divide it into a training set and a test set;

[0007] Step S2: Extract Log-Mel spectrogram features and Gammatone spectrogram features from the audio data samples in the dataset, and perform fusion processing on the Log-Mel spectrogram and the Gammatone spectrogram to obtain L-Gtg features;

[0008] Step S3: Based on the FastViT network, construct an improved DFHL-FastViT network. Specifically: Add a DF dual-channel feature fusion module; Modify the multi-head attention module (MHSA) and replace it with a HiLo attention module to obtain the DFHL-FastViT network;

[0009] Input the training set in S2 into the improved DFHL-FastViT network in S3 for model training to obtain a trained DFHL-FastViT network;

[0010] Step S4: Input the test set into the trained DFHL-FastViT network in S3 to identify the classification of indoor abnormal sounds.

[0011] The beneficial effects of the present invention are:

[0012] The purpose of the present invention is to address the problem that the masking effect generated under high-level background noise conditions in the indoor environment affects the recognition of abnormal sounds. This paper proposes an indoor abnormal sound recognition method based on L-Gtg features and an improved FastViT model. First, extract Log-Mel and Gammatonegram features from the original audio to form new L-Gtg features, increasing the diverse representation of features. Then, propose a DFHL-FastViT audio classification model, in which a dual-channel feature fusion module is designed to dynamically allocate weights to different feature channels of L-Gtg for fusion, and the HiLo attention module is introduced to improve the MHSA attention module in FastViT to capture the global relationship and local sudden changes in the time-frequency features, while reducing the computational cost and the overall complexity of the model. The number of parameters of the improved model is reduced by 1.11M, a total decrease of 5.39%. And the experimental results show that the method proposed in the present invention, even in a relatively high indoor background noise environment, not only has a robust identification ability with a high correct rate but also has a certain anti-noise property. Description of the Drawings

[0013] Figure 1 is the flowchart of the detection method of the present invention;

[0014] Figure 2 is the structural diagram of the DFHL-FastViT network in the present invention;

[0015] Figure 3 is the structural diagram of the DF dual-channel feature fusion module in the present invention;

[0016] Figure 4 It is the structural diagram of the HiLo attention module in the present invention; Detailed implementation manners

[0017] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art. The specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] Embodiment

[0019] An indoor abnormal sound recognition method based on L-Gtg features and an improved FastViT model, as Figure 1 shown, includes the following steps:

[0020] S1. From the audio samples in the public datasets (ESC50, Acoustic Event Dataset (AED)), select 10 types of abnormal sound categories that meet the indoor scene to form an indoor abnormal sound dataset. And select restaurant noise, fan sound, and kitchen frying sound from the Noise-x92 dataset and the Disco dataset as background noises, and mix them with the abnormal sounds according to different signal-to-noise ratio (SNR) conditions (-20dB, -15dB, -10dB, 0dB) to generate a noisy indoor abnormal sound dataset. The total number of samples in the dataset is 5538 audio samples. Then divide it into a training set and a test set. The time length of each audio sample in the training set and the test set is 5 seconds, and the sampled audio is 16kHz.

[0021] S2. Extract Log-Mel spectrogram features and Gammatone spectrogram features from the audio data samples in the dataset; among them, the frame size of the Log-Mel spectrogram is 4096, the number of Mel filters is 128, the frame size of the Gammatone spectrogram is 4096, and the number of Mel filters is 128.

[0022] S3. Based on the FastViT network, construct an improved DFHL-FastViT network, specifically: add a DF dual-channel feature fusion module, and then replace the MASH multi-head self-attention module with a HiLo module to obtain the improved network DFHL-FastViT.

[0023] Input the training set in S2 into the DFHL-FastViT network for model training to obtain a trained network model.

[0024] S4. Input the test set into the DFHL-FastViT model trained in S3 to identify the classification of indoor abnormal sounds.

[0025] FastViT combines the advantages of convolution and Transformer. During the training phase, complex modules are adopted to enhance the model's expressive ability and feature learning ability. While in the inference phase, through the structural reparameterization technique, the complex structure in the training phase is simplified, reducing the memory access cost, thereby improving the inference speed and efficiency. This design not only endows the model with powerful feature extraction capabilities, enabling it to accurately capture complex time-frequency information in an indoor noise environment, but also significantly improves the inference speed, making it very suitable for application scenarios of real-time processing of abnormal sound events.

[0026] On this basis, the present invention proposes the DFHL-FastViT model, and its main structure is as Figure 2 shown. By designing a dual-channel feature fusion learning module and drawing on the ideas of multi-scale features and gated units, it dynamically assigns weights and fuses different channels of the L-Gtg features, enhancing the model's learning ability for multi-modal features.

[0027] The structure of the dual-channel feature fusion learning module is as Figure 3 shown. The dual-channel feature fusion learning module can dynamically adjust the relative importance between different features by learning weights. Specifically, first, global average pooling (Global Average Pooling, GAP) is performed on the dual-channel features to obtain the global information of the features, and the Sigmoid function is used to map the weight values to the interval [0, 1] to obtain the preliminary weights. Then, the Softmax function is used to normalize the weights so that the weights are reasonably distributed between the two different features. Finally, based on these learned weights, weighted fusion is completed to generate the final fused feature representation. This feature fusion strategy enables the model to better understand the multi-scale features of audio, can make more effective use of these two features during the training process, and thus improves the task performance. Its calculation formula is as follows:

[0028] (1)

[0029] Where, and are the input features; is the Sigmoid function; is the global average pooling layer; is the result after global average pooling and Sigmoid operation and concatenation.

[0030] (2)

[0031] Where, Softmax is the Softmax function; W is the weight vector, which can be expressed as:

[0032] (3)

[0033] Among them, and respectively correspond to the weights of features and .

[0034] (4)

[0035] Finally, based on the weight vectors and the input features are weighted and fused to obtain the dual-channel fusion feature .

[0036] When the multi-head self-attention mechanism MHSA adopted by it processes audio features, it requires the model to globally focus on all feature blocks. In the case of noise interference in a complex indoor environment, it may lead to difficulties in highlighting abnormal sound information masked by background noise in the spectrogram, as well as problems such as a large number of model parameters and high computational overhead. The present invention introduces the HiLo attention module to modify the MHSA module in FastViT, enabling the model to effectively capture the global relationship and local sudden changes in time-frequency features, reducing the computational cost and complexity, and making it more suitable for the abnormal sound recognition task in a complex indoor noise environment. The HiLo attention module is as shown in Figure 4 . The HiLo attention module (HiLo Attention) inputs the features into two groups of attention, namely the high-frequency attention and the low-frequency attention branches. It focuses on key abnormal acoustic features by capturing the global relationship and local sudden changes in time-frequency features.

[0037] In the embodiment, the indoor abnormal sound recognition method proposed by the present invention is adopted. While the number of model parameters is reduced by 1.11M, a total decrease of 5.39%, under the background noise condition with a signal-to-noise ratio (SNR) of -15 dB, an accuracy rate of 93.83% is achieved. It is confirmed that the improved algorithm not only has a robust identification ability with a high correct rate even in a high indoor background noise environment, but also has a certain anti-noise property, which has a positive effect on the indoor abnormal sound recognition task with a high noise level.

[0038] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. An indoor abnormal sound recognition method based on L-Gtg features and an improved FastViT model; characterized in that, It includes the following steps: S1: Select 10 abnormal sound categories that meet the indoor scene from the audio samples in the public datasets (ESC50, Acoustic Event Dataset (AED)) to form an indoor abnormal sound dataset; then divide the data to generate a training set and a test set; S2: Extract Log-Mel spectrogram features and Gammatone spectrogram features from the audio data samples in the dataset, and perform fusion processing on the Log-Mel spectrogram and Gammatone spectrogram to obtain L-Gtg features; S3: Based on the FastViT network, construct an improved DFHL-FastViT network. Specifically: add a DF dual-channel feature fusion module; modify the multi-head attention module (Multi-Head Self-Attention, MHSA) and replace it with a HiLo attention module to obtain the DFHL-FastViT network; Input the training set in S2 into the improved DFHL-FastViT network in S3 for model training to obtain a trained DFHL-FastViT network; S4: Input the test set into the trained DFHL-FastViT network in S3 to identify the classification of indoor abnormal sounds.

2. The indoor abnormal sound recognition method according to claim 1, wherein: The dataset used in step S1 contains 10 abnormal sound categories that meet the indoor scene, specifically including "fire sound", "scream", "baby cry", "broken glass sound", "knocking sound", "gunshot sound", "talking sound", "water drop sound", "alarm sound" and "dog barking sound"; at the same time, "restaurant noise", "fan sound" and "kitchen frying sound" are selected from the Noise-x92 and Disco datasets as background noises; according to different signal-to-noise ratio (SNR) conditions (-20dB, -15dB, -10dB, 0dB respectively), these background noises are mixed with abnormal sounds to generate a noisy indoor abnormal sound dataset; this dataset contains a total of 5538 audio samples, and the training set and test set are divided in a ratio of 8:

2.

3. The indoor abnormal sound recognition method according to claim 1, characterized in that: The frame size of the Log-Mel spectrogram in S2 is 4096, the number of Mel filters is 128, the frame size of the Gammatone spectrogram is 4096, and the number of filters is 128.

4. The indoor abnormal sound recognition method according to claim 1, wherein: In step S3, it is based on improving the FastViT network; first, add a DF dual-channel feature fusion module; Then modify the multi-head attention module (Multi-Head Self-Attention, MHSA) and replace it with a HiLo attention module to obtain the DFHL-FastViT network; the calculation formula of the DF dual-channel feature fusion module is as follows: (1) Among them, and are input features; is the Sigmoid function; is the global average pooling layer; is the result concatenated after global average pooling and Sigmoid operations; (2) Among them, Softmax is the Softmax function; W is the weight vector, which can be expressed as: (3) Among them, and correspond to the weights of features and respectively; (4) Finally, based on the weight vectors and the input features are weighted and fused to obtain dual-channel fused features .