A Sound Event Localization and Recognition Method Based on a Residual Module with a Fusion Channel Attention Mechanism

By adopting the residual module based on the fusion channel attention mechanism in sound event positioning and detection, the problem of insufficient accuracy of sound event positioning and detection in complex environments is solved, and efficient multi-sound source positioning and separation are achieved, and the stability and performance of the model are improved.

CN116631386BActive Publication Date: 2025-06-13GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310245365.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2025-06-13
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

Existing sound event positioning and detection methods have problems with insufficient accuracy and robustness in complex environments, especially under the influence of factors such as environmental noise, signal attenuation, multi-path propagation and interference, it is difficult to effectively identify and locate multiple sound sources.

Method used

The sound event positioning and recognition method of the residual module based on the fusion channel attention mechanism is adopted. This method improves the model's feature extraction ability and the fusion of global spatial information through the SE residual block and the extrusion and excitation network module, and simultaneously detects and positions sound events through a multi-task learning framework, and optimizes the loss function to improve the generalization ability and stability of the model.

Benefits of technology

It realizes accurate identification and positioning of multiple sound sources in complex environments, reduces algorithm complexity and calculation amount, improves the performance and stability of the model, and can significantly improve the accuracy of sound event positioning and detection without increasing the complexity of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631386B_ABST
    Figure CN116631386B_ABST
Patent Text Reader

Abstract

The present invention provides a method for sound event localization and recognition based on a residual module with a fusion channel attention mechanism. This method uses an SE residual block to improve the feature extraction ability of the network and the fusion of spatial information. At the same time, it can perform sound event detection and sound event localization simultaneously, reducing the algorithm complexity and computational amount. The loss functions of the sound event detection and sound event localization tasks are optimized using a joint training method, improving the generalization ability and stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound signal processing, and particularly to a method for sound event localization and recognition based on a residual module integrating a channel attention mechanism. Background Art

[0002] Sound event localization and detection refer to identifying and locating specific sound events from multiple audio signal sources, such as vehicle horns, human voices, etc. This research has important practical applications, such as in the fields of intelligent audio monitoring, autonomous driving, and security monitoring. In practical applications, sound event localization and detection are affected by various factors, such as environmental noise, signal attenuation, multipath propagation, and interference. Therefore, researchers need to develop new algorithms and technologies to address these challenges. Traditional methods include using sensor arrays and signal processing techniques to locate sound sources, but these methods have some limitations in practical applications, such as high costs, the need for complex equipment, and a large amount of computing resources.

[0003] In recent years, the rapid development of deep learning technology has brought new opportunities for solving sound event localization and detection. Deep learning technology can automatically extract features from data and can learn complex non-linear models, with high accuracy and robustness. Therefore, many researchers have begun to explore applying deep learning technology to sound event localization and detection. However, existing network structures have some problems, such as incomplete feature extraction of data and lack of attention to the spatial information features between data channels. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method for sound event localization and recognition based on a residual module integrating a channel attention mechanism. This method uses an SE residual block to improve the network's feature extraction ability and the fusion of spatial information. At the same time, it can achieve both sound event detection and sound event localization, reducing the algorithm complexity and computational amount. The loss functions of the sound event detection and sound event localization tasks are optimized using a joint training method, improving the generalization ability and stability of the model.

[0005] The present invention is achieved through the following technical solutions:

[0006] A method for sound event localization and recognition based on a residual module integrating a channel attention mechanism, characterized by comprising the following steps:

[0007] Step 1: Collection and preprocessing of a sound event dataset. Collect a multi-channel audio dataset using first-order Ambisonics (FOA) format signals, and then preprocess the data to obtain multi-channel logarithmic Mel spectrograms and FOA intensity vectors;

[0008] Step 2: Feature extraction. Input the logarithmic Mel spectrogram and intensity vector obtained in Step 1 into the residual module network with a fused channel attention mechanism to extract the required features;

[0009] Step 3: Audio event detection (SED): Use a fully connected neural network to classify the features obtained in Step 2 at each moment. At each time step, the SED task outputs a binary classification label indicating whether there is a sound event at that time step to determine whether there is an audio event.

[0010] Step 4: Audio event localization (SEL): Use another regression task for the features obtained in Step 2 at each moment. At each time step, the SEL task outputs a quadruple representing the location and duration of the sound event in three-dimensional space.

[0011] Step 5: Multi-task learning: Divide the audio data into a training set, a validation set, and a test set. Build a time-domain convolutional neural network to train the audio data, combine the loss functions of the SED and SEL tasks, and use the method of joint training for optimization.

[0012] Step 6: Evaluate the model using predefined evaluation metrics and compare it with other methods to determine whether its performance is excellent enough.

[0013] As an optimization of the technical solution, in Step 1, for the preprocessing method of audio data, the collected data is saved in the four-channel FOA format with a sampling frequency of 24KHz. Use a 1024-point FFT, a 40-millisecond Hanning window, and a 20-millisecond hop length to calculate the four-channel spectrogram, and extract the logarithmic Mel spectrogram of 64 Mel bands from the spectrogram; calculate the spatial features of the acoustic intensity vector of each STFT frequency band from the FOA spectrogram and aggregate them into a similar number of Mel bands to provide input for the network.

[0014] As an optimization of the technical solution, in Step 3, for the sound event localization and recognition decoder, we use a bilinear gated recurrent unit, followed by layer normalization and hyperbolic tangent activation. Then, obtain the sound event localization and recognition output by applying two fully connected layers and hyperbolic tangent activation. Finally, for each moment, calculate the location and duration of each event according to the predicted class probability.

[0015] As an optimization of the technical solution, in step two, an SE residual module is used. The residual block enhances the receptive field of the network, further explores the multi-scale expression ability of the convolutional neural network at a finer granularity level, and significantly improves the feature extraction ability of the network; the squeeze-and-excitation module squeezes the time-domain and frequency-domain coordinates of each channel of the sound into a scalar value, then performs excitation on it, and finally performs a re-weighting operation with the corresponding channel feature map to achieve global spatial information fusion.

[0016] As an optimization of the technical solution, in step five, the model uses a multi-task learning framework to simultaneously learn the SED and SEL tasks. Specifically, we use two different output branches, one for predicting the SED task and the other for predicting the SEL task. The two branches share the same input features and learn the feature representation through shared convolutional layers. At the same time, we use different classifiers to handle the classification of the SED and SEL tasks. During the training process, we adopt a multi-task loss function to simultaneously optimize the performance of the two tasks, so that the model can handle the SED and SEL tasks simultaneously.

[0017] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0018] 1. The present invention provides a method for sound event localization and recognition based on a residual module with a fused channel attention mechanism, which is used to accurately identify and locate multiple sound sources in a complex environment. This method uses a network with a residual module with a fused channel attention mechanism to process the input audio signal, and technologies such as residual blocks and squeeze-and-excitation network modules are adopted to improve the feature extraction ability of the model and the fusion of global spatial information. This method can not only identify the categories of sound sources but also estimate their positions in space, thus realizing the localization and separation of multiple sound sources.

[0019] 2. The method of the present invention does not need to use complex feature engineering to extract useful information from audio signals, but uses a network with a residual module with a fused channel attention mechanism to automatically learn the features in audio signals, which can improve the performance of the model without increasing the model complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The following are the main drawings of this method.

[0021] Figure 1 It is a schematic diagram of the data processing flow provided by the present invention.

[0022] Figure 2 It is a schematic diagram of the overall structure of the residual module network with a fused channel attention mechanism of the present invention.

[0023] Figure 3 It is a structural diagram of the residual block of the present invention. DETAILED DESCRIPTION

[0024] The present invention is further described in detail below by way of examples. These examples are only used to illustrate the present invention and do not limit the protection scope of the present invention.

[0025] like Figure 1 As shown, the present invention provides a sound event localization and recognition method based on a residual module of a fusion channel attention mechanism, comprising the following steps:

[0026] The audio dataset was collected using Kinect2.0's four-channel microphone array, which was mounted on a tripod 1.2 meters above the ground. Fourteen sounds were collected indoors, including alarm clock sounds, baby crying, door knocking, dog barking, walking sounds, piano sounds, etc., and the type and location information of the sounds were annotated.

[0027] (1) Preprocessing of audio data: The collected data is saved in a four-channel FOA format with a sampling frequency of 24 kHz. A four-channel spectrogram is calculated using a 1024-point FFT, a 40-ms Hanning window, and a 20-ms jump length. The logarithmic Mel spectrogram of 64 mel bands is extracted from the spectrogram. The spatial features of the acoustic intensity vector of each STFT frequency band are calculated from the FOA spectrogram and aggregated into a similar number of Mel bands to provide input for the network.

[0028] (2) Feature extraction of audio data: The spectrum graph is input into the network model for feature extraction. The SE residual module in the model expands the network's receptive field, deepens the network structure and optimizes the network layer. The channel attention mechanism compresses and excites the channel features and weights them to obtain a high-level feature graph. After that, it enters the fully connected network to regress and classify the features for judgment, detection and positioning.

[0029] (3) Multi-task learning: The audio data is divided into a training set, a validation set, and a test set, which are divided into 6 cross-validation parts, with 100 recordings in each part. 400 recordings are used for training, 100 recordings for validation, and 100 recordings for testing. A time-domain convolutional neural network is built to train the audio data, and the loss functions of the SED and SEL tasks are combined and optimized using a joint training method;

[0030] (4) Prediction results and model evaluation: The model is evaluated using pre-defined evaluation metrics and compared with other methods to determine whether its performance is good enough;

[0031] (5) Evaluation metrics: For audio event detection, the F score and error rate (ER) calculated in units of 1 second are used. For audio event localization estimation, two frame-by-frame metrics, audio event localization error (DE) and frame recall (FR), are used.

[0032] Table 1 Prediction Results of Different Models for Sound Event Localization and Detection

[0033]

[0034] Table 1 shows the quantitative comparison between the SE-ResNet and SELD proposed by the present invention and ResNet, and there are significant performance improvements on the collected data sets; the network structure proposed by the present invention enhances the expression of key channel features and spatial features, and at the same time, the deep network avoids the problems of network degradation and gradient disappearance, and optimizes the network structure.

[0035] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for sound event localization and recognition based on a residual module with a fused channel attention mechanism, characterized in that, it includes the following steps: Step 1: Acquisition and preprocessing of the sound event dataset. Collect a multi-channel audio dataset, use the first-order Ambisonics FOA format signal, and then preprocess the data to obtain the multi-channel log Mel spectrogram and FOA intensity vector; Step 2: Feature extraction. Input the log Mel spectrogram and intensity vector obtained in Step 1 into the residual module network with a fused channel attention mechanism to extract the required features; Step 3: Audio event detection SED: Use a fully connected neural network to classify the features obtained in Step 2 at each moment. At each time step, the SED task outputs a binary classification label indicating whether there is a sound event at that time step to determine whether there is an audio event; Step 4: Audio event localization SEL: Use another regression task for the features obtained in Step 2 at each moment. At each time step, the SEL task outputs a quadruple indicating the position and duration of the sound event in three-dimensional space; Step 5: Multi-task learning: Divide the audio data into a training set, a validation set, and a test set, build a residual with a fused channel attention mechanism to train the audio data, combine the loss functions of the SED and SEL tasks, and use the method of joint training for optimization; Step 6: Use predefined evaluation metrics to evaluate the model and compare it with other methods to determine whether its performance is excellent enough.

2. The method for sound event localization and recognition based on a residual module with a fused channel attention mechanism according to claim 1, characterized in that, in Step 1, the preprocessing method of the audio data is to save the collected data in the four-channel FOA format, with a sampling frequency of 24KHz, calculate the four-channel spectrogram using a 1024-point FFT, a 40-millisecond Hanning window, and a 20-millisecond hop length, and extract the log Mel spectrogram of 64 mel bands from the spectrogram; calculate the spatial features of the acoustic intensity vector of each STFT frequency band from the FOA spectrogram and aggregate them into a similar number of Mel bands to provide input for the network.

3. The method for sound event localization and recognition based on a residual module with a fused channel attention mechanism according to claim 1, characterized in that, in Step 3, for the sound event localization and recognition decoder, we adopt a bilinear gated recurrent unit, followed by layer normalization and hyperbolic tangent activation; then, obtain the sound event localization and recognition output by applying two fully connected layers and hyperbolic tangent activation; finally, for each moment, calculate the position and duration of each event according to the predicted class probability.

4. The method for sound event localization and recognition based on a residual module with a fused channel attention mechanism according to claim 1, characterized in that, In step two, the SE residual module is used; the residual block enhances the receptive field of the network, further explores the multi-scale representation ability of the convolutional neural network at a finer granularity level, and greatly improves the feature extraction ability of the network; the squeeze-and-excitation module squeezes the time-domain and frequency-domain coordinates of each channel of the sound into a scalar value, then performs excitation on it, and finally performs a re-weighting operation with the corresponding channel feature map to achieve global spatial information fusion.

5. The method for sound event localization and recognition based on a residual module with a fusion channel attention mechanism according to claim 1, wherein, in step five, the model uses a multi-task learning framework to simultaneously learn the SED and SEL tasks; specifically, we use two different output branches, one for predicting the SED task and the other for predicting the SEL task; the two branches share the same input features and learn feature representations through a shared convolutional layer; at the same time, we use different classifiers to handle the classification of the SED and SEL tasks; during the training process, we adopt a multi-task loss function to simultaneously optimize the performance of the two tasks, so that the model can handle the SED and SEL tasks simultaneously.

Citation Information

Patent Citations

  • Specific sound event retrieval and positioning method based on sequence classification

    CN111161715A

  • Sound event detection and positioning method based on deep learning

    CN113921034A