Audio recognition method based on convolutional neural network with one-dimensional attention mechanism
By using a one-dimensional attention mechanism convolution neural network in audio signal processing and integrating channel and time attention mechanism, the high demand for data and computing resources and information loss problems in 2DCNN technology, as well as the poor performance of one-dimensional convolutional neural network, achieving efficient and accurate audio signal recognition.
Patent Information
- Application Number
- CN202111611392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The existing 2DCNN technology requires a large amount of data and computing resources in audio signal processing, and may lead to information loss; while the model based on one-dimensional convolutional neural networks is poor in performance and poor in classification effect.
The audio recognition method based on the one-dimensional attention mechanism convolutional neural network is adopted to improve feature learning and classification performance by integrating attention mechanism modules, including channel attention mechanism and time attention mechanism, in the one-dimensional convolutional neural network.
Through the one-dimensional attention mechanism convolutional neural network, efficient, accurate and stable audio signal classification recognition can be achieved with less data and smaller calculation amount, improving the accuracy of audio signal recognition.
Smart Images

Figure CN114187923B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio signal processing, and in particular to an audio recognition method based on a one-dimensional attention mechanism convolutional neural network. Background Art
[0002] In recent years, deep learning has flourished in various fields, especially convolutional neural networks, which have had a significant impact on the processing of audio and music, such as automatic music tagging, speaker recognition, and environmental sound classification. The commonly used method is 2DCNN (Chinese meaning: two-dimensional convolution), which first converts the audio signal into a two-dimensional spectrogram using time-frequency analysis (such as short-time Fourier transform), and then uses the CNN (Chinese meaning: convolutional neural network, English full name Convolutional Neural Networks) architecture commonly used in image recognition (such as AlexNet, VGGNet) for recognition and analysis. However, since 2DCNN has more parameters, in practical applications, a large amount of data is required to obtain better generalization performance. In order to solve the technical problems of 2DCNN, 1DCNN (Chinese meaning: one-dimensional convolution) technology that directly processes the original audio signal has been produced in recent years. This technology has achieved good performance in many low-resource databases, such as early arrhythmia detection in electrocardiogram beats, structural health monitoring, and structural damage detection.
[0003] Although, compared with the traditional 2DCNN technology, the 1DCNN technology learns directly from the original waveform of the signal and can utilize the fine time structure of the signal; and does not require the additional process of converting the audio signal from 1D to 2D, that is, it does not need to convert the audio signal into a two-dimensional spectrogram using time-frequency analysis (such as short-time Fourier transform), thus avoiding the problem of irreversible loss of useful information in the process of converting from 1D to 2D; at the same time, the 1DCNN technology directly learns from the one-dimensional original waveform of the signal, which effectively reduces the amount of calculation compared to the 2DCNN technology. However, the model trained based on the one-dimensional convolutional neural network is limited in complexity, so the network performance is poor and the classification effect is poor. Summary of the invention
[0004] The present invention aims to provide an audio recognition method based on a one-dimensional attention mechanism convolutional neural network.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for audio recognition based on a one-dimensional attention mechanism convolutional neural network, comprising the following steps:
[0007] S1, construct the original dataset;
[0008] Downloading publicly available audio signal data to construct the original data set;
[0009] S2, data preprocessing;
[0010] Use Python code to perform frame processing on the audio signal in the original data set;
[0011] S3, establish a one-dimensional attention mechanism convolutional neural network;
[0012] The one-dimensional attention mechanism convolutional neural network integrates the attention mechanism module into the one-dimensional convolutional neural network.
[0013] The one-dimensional convolutional neural network includes an input layer, a convolution layer, an activation function, a pooling layer, a fully connected layer and a Softmax layer;
[0014] The attention mechanism module includes a channel attention mechanism module and a time attention mechanism module;
[0015] S4, training a one-dimensional attention mechanism convolutional neural network;
[0016] The audio signals in the original data set are divided into independent data sets according to the ratio of training set: validation set: test set = 8:1:1; the audio signals in the training set are input into a one-dimensional attention mechanism convolutional neural network, and the Ranger neural network optimizer, mean square logarithmic loss function, and cosine annealing learning rate are used to complete the training of the one-dimensional attention mechanism convolutional neural network to obtain multiple models; and the validation set and the test set are used to complete the verification and testing of the model;
[0017] S5, integrates multiple models through ensemble learning to form the final audio recognition model;
[0018] Preferably, in step S3, the input layer is used to receive the audio signal; the activation function is ReLU; the pooling layer adopts one-dimensional pooling with global average; the convolution layer is a one-dimensional convolution that can reduce the amount of calculation and speed up the calculation; and the Softmax layer is used to classify the training results.
[0019] Preferably, in step S3, the channel attention mechanism module aggregates the feature information of the feature maps output by each channel of the convolutional layer, recalibrates the weights of the feature maps of each channel, distinguishes the importance of the feature maps of each channel by different weights, obtains the channels that the one-dimensional convolutional neural network needs to pay special attention to, suppresses the channels with less effect, and adaptively changes the weights of the feature maps of each channel.
[0020] Preferably, in step S3, the temporal attention mechanism module aggregates feature information of feature maps output by each channel of the convolutional layer to locate the time signal segment related to the audio signal.
[0021] Preferably, the step S5 specifically includes selecting models of ten different nodes from the models obtained in the step S4 that converge to multiple different minimum values, saving the model parameter weights, and performing model integration to form the final audio recognition model to obtain a higher audio recognition accuracy.
[0022] Preferably, in step S2, the frame processing is performed according to a frame length of 1 s, an overlap ratio of 0.75, and a sampling rate of 8000 Hz.
[0023] The advantage of the present invention is that it uses a one-dimensional attention mechanism convolutional neural network to recognize audio signals, which not only overcomes the problems of large original data requirements, large amount of calculation, and loss of some useful information of two-dimensional convolutional neural networks, but also overcomes the problems of poor network performance and poor classification effect of one-dimensional convolutional neural networks. By forming a more efficient and powerful audio signal recognition model with a one-dimensional attention mechanism convolutional neural network, it is possible to achieve efficient, accurate and stable classification and recognition of audio signals with less original data and less calculation. Experiments have shown that the accuracy of audio signal recognition on the public data set Urbansound8K can reach about 93%, and can reach 94% with the addition of ensemble learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the method of the present invention.
[0025] Figure 2 It is a schematic diagram of the channel attention mechanism structure of the method described in the present invention.
[0026] Figure 3 Schematic diagram of the temporal attention mechanism structure of the method described in the present invention.
[0027] Figure 4 It is a schematic diagram of the one-dimensional convolutional neural network structure of the method described in the present invention. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] like Figure 1As shown, the audio recognition method based on a one-dimensional attention mechanism convolutional neural network of the present invention comprises the following steps:
[0030] S1, construct the original dataset;
[0031] In this example, Urbansound8K, a public environmental sound dataset, is downloaded as the audio signal to be processed to construct the original dataset. Urbansound8K contains 8732 labeled data, each of which is no longer than 4 seconds, including air conditioner sound, car horn, children playing, dog barking, drilling, engine idling, gunshot, jackhammer, police siren, and street music, a total of ten categories.
[0032] S2, data preprocessing;
[0033] For the audio signals in Urbansound8K, we use Python code to perform frame processing according to the frame length of 1s, overlap ratio of 0.75, and sampling rate of 8000 Hz. The number of audio data after frame processing increases from 8732 to about 30,000, which greatly expands the data volume of the original data set and meets the usage requirements.
[0034] S3, establish a one-dimensional attention mechanism convolutional neural network;
[0035] like Figure 2 As shown, the one-dimensional attention mechanism convolutional neural network integrates the attention mechanism module into the one-dimensional convolutional neural network to improve the feature learning ability and classification performance of the one-dimensional convolutional neural network.
[0036] Among them, the one-dimensional convolutional neural network includes an input layer, a convolution layer, an activation function, a pooling layer, a fully connected layer and a Softmax layer; the input layer is used to receive the audio signal; the activation function is ReLU; the pooling layer adopts one-dimensional pooling with global average; the convolution layer is a one-dimensional convolution that can reduce the amount of calculation and speed up the calculation speed; the Softmax layer is used to classify the training results.
[0037] The attention mechanism module includes but is not limited to a channel attention mechanism module and a time attention mechanism module; it is used to embed in convolutional neural networks of different depths, adaptively optimize feature mapping, and accumulate this advantage in the entire convolutional neural network by adding multiple layers of attention modules. In addition, the attention mechanism module can learn the relationship between the target output and the time input, explore and find the most relevant feature maps and time signal segments for audio signals of different classifications, thereby enhancing relevant features, suppressing irrelevant features, and ultimately improving the classification effect of the convolutional neural network.
[0038] Among them, Figure 3As shown in the figure, the channel attention mechanism module aggregates the feature information of the feature maps output by each channel of the convolutional layer, recalibrates the weights of the feature maps of each channel, distinguishes the importance of the feature maps of each channel by different weights, obtains the channels that the convolutional neural network needs to pay attention to, suppresses the channels with less effect, and adaptively changes the weights of the feature maps of each channel.
[0039] like Figure 4 As shown in the figure, the temporal attention mechanism module aggregates the feature information of the feature maps output by each channel of the convolutional layer to locate the time signal segment related to the audio signal, thereby optimizing the feature response of the convolutional neural network, enhancing the time features related to the audio signal, ignoring the time features not related to the audio signal, improving the efficiency of the convolutional neural network, and speeding up the training of the convolutional neural network model.
[0040] Specifically, for multi-channel input Y=yi, global average pooling is used to compress the global temporal information into one channel to generate a channel statistical vector z, where the i-th element z is shown in Formula 1:
[0041] (1)
[0042] Then CAM (Chinese interpretation is channel attention mechanism) adopts a simple gating mechanism to fully capture the channel-wise dependencies and generate channel recalibration vectors. As shown in Formula 2:
[0043] (2)
[0044] in is the ReLU activation function, , Represents a convolution with a channel number of 1 and a convolution kernel size of 1x1. is the Sigmoid function, which compresses information into the range of [0,1]. The value of represents the importance of the i-th channel. The channel recalibration vector M is used to recalibrate the feature Y as shown in Formula 3:
[0045] (3)
[0046] Finally, the obtained feature M fully considers the guidance of global information and can effectively highlight more discriminative feature information. The idea of residual learning is used and residual connections are introduced to optimize and reacquire the original information. The final output of CAM is: YCAM=Y+M.
[0047] The input features are ,in , represents the jth time signal position, j=1,2, … W. TAM obtains the feature map s of the input Y in the time domain signal through a 1x1 convolutional layer as shown in Formula 4:
[0048] (4)
[0049] The 1x1 convolution operation can aggregate the features of all channels, and then obtain the time weight vector through the Sigmoid function. As shown in Formula 5:
[0050] (5)
[0051] The time weight vector can indicate the importance of the time series point. Multiplying the time weight vector and the feature map to obtain the recalibrated feature map N is shown in Formula 6:
[0052] (6)
[0053] Before recalibration, TAM (Chinese translation is Temporal Attention Mechanism) uses convolutional layers to encode feature information between local temporal signal segments to prevent over-focusing on related temporal signal segments. Similar to CAM, residual connections are introduced to prevent the reduction of feature response values in TAM. Finally, TAM output is YTAM=Y+N.
[0054] S4, training a one-dimensional attention mechanism convolutional neural network;
[0055] The original data set of about 30,000 audio signals is divided into three independent data sets according to the ratio of training set: validation set: test set = 8:1:1. Among them, there are about 24,000 audio signals in the training set, about 3,000 audio signals in the validation set, and about 3,000 audio signals in the test set. After that, the audio signals in the training set are input into the one-dimensional attention mechanism convolutional neural network for training. The Ranger neural network optimizer is selected, and the learning rate is set to 0.001. The loss function is set to the mean square logarithm loss function. The cosine annealing learning rate is used. Only one training is required to obtain a model that converges to multiple different minimum values, completing the training of the one-dimensional attention mechanism convolutional neural network. At the same time, the model is verified and tested through the validation set and the test set.
[0056] S5, integrates multiple models through ensemble learning to form the final audio recognition model.
[0057] Select models with ten different nodes from the models obtained in step S4 that converge at multiple different minimum values, save the model parameter weights, perform model integration, and form the final audio recognition model, which can achieve higher audio recognition accuracy.
Claims
1. A method for audio recognition based on a one-dimensional attention mechanism convolutional neural network. Features: The following steps are involved: S1, construct the original dataset; Downloading publicly available audio signal data to construct the original data set; S2, data preprocessing; Use Python code to perform frame processing on the audio signal in the original data set; S3, establish a one-dimensional attention mechanism convolutional neural network; The one-dimensional attention mechanism convolutional neural network integrates the attention mechanism module into the one-dimensional convolutional neural network; The one-dimensional convolutional neural network includes an input layer, a convolution layer, an activation function, a pooling layer, a fully connected layer and a Softmax layer; The attention mechanism module includes a channel attention mechanism module and a time attention mechanism module; the channel attention mechanism module aggregates the feature information of the feature map output by each channel of the convolution layer, recalibrates the weight of the feature map of each channel, distinguishes the importance of the feature map of each channel by different weights, obtains the channels with high importance in the one-dimensional convolutional neural network, suppresses the channels with low importance, and adaptively changes the weight of the feature map of each channel; The temporal attention mechanism module aggregates the feature information of the feature graph output by each channel of the convolutional layer to locate the time signal segment related to the audio signal; S4, training a one-dimensional attention mechanism convolutional neural network; The audio signals in the original data set are divided into independent data sets according to the ratio of training set: validation set: test set = 8:1:1; the audio signals in the training set are input into the one-dimensional attention mechanism convolutional neural network, and the Ranger neural network optimizer, mean square logarithmic loss function, and cosine annealing learning rate are used to complete the training of the one-dimensional attention mechanism convolutional neural network to obtain multiple models; and the validation set and the test set are used to complete the verification and testing of the model; S5, integrates multiple models through ensemble learning to form the final audio recognition model.
2. The method for audio recognition based on a one-dimensional attention mechanism convolutional neural network according to claim 1, Features: In step S3, the input layer is used to receive the audio signal; the activation function is ReLU; the pooling layer adopts one-dimensional pooling with global average; the convolution layer is a one-dimensional convolution that can reduce the amount of calculation and speed up the calculation; the Softmax layer is used to classify the training results.
3. The audio recognition method based on a one-dimensional attention mechanism convolutional neural network according to claim 1, Features: The step S5 specifically includes selecting models of ten different nodes from the models obtained in step S4 that converge to multiple different minimum values, saving the model parameter weights, and performing model integration to form the final audio recognition model to obtain a higher audio recognition accuracy.
4. The audio recognition method based on a one-dimensional attention mechanism convolutional neural network according to claim 1, Features: In step S2, the frame processing is performed according to a frame length of 1 s, an overlap ratio of 0.75, and a sampling rate of 8000 Hz.
Citation Information
Patent Citations
Convolutional neural network optimization method based on attention
CN108875592A
Human body posture recognition method based on time and channel double attention
CN111860188A