A Sound Source Localization Method Based on a Four-Microphone Array and Deep Learning

By using a combination of four microphone arrays and deep learning in the sound source positioning algorithm, the convolutional recurrent neural network is improved by using residual blocks and attention mechanisms, the problem of insufficient generalization ability of the sound source positioning algorithm in unknown environments in the existing technology is solved, and higher positioning accuracy and robustness are achieved.

CN116106827BActive Publication Date: 2025-06-13HAINACORD (HUBEI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211727267.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-06-13
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

In the prior art, the deep learning-based sound source positioning algorithm has low generalization capabilities for unknown environments (noise and reverb), and its performance needs to be further improved.

Method used

The sound source positioning method based on four microphone arrays and deep learning is adopted to process the original sound source data through short-time Fourier transform, extract STFT phase characteristics, and use residual blocks and attention mechanisms to improve the convolutional recurrent neural network to perform sound source positioning.

Benefits of technology

The generalization ability of the sound source positioning algorithm in unknown environments is improved, the ability to screen input features is enhanced, and the robustness of the model and the utilization efficiency of hardware resources is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116106827B_ABST
    Figure CN116106827B_ABST
Patent Text Reader

Abstract

The present invention discloses a sound source localization method based on a four-microphone array and deep learning. The sound source signal is collected through a tetrahedral microphone array equipped with four microphones to obtain the original sound source audio information. The original sound source data is subjected to short-time Fourier transform to convert it into a phase spectrum, and the phase spectrum is input into a neural network for training. The trained model is used to predict the sound source angle information. The beneficial effects of the present invention are as follows: On the basis of the traditional convolutional recurrent neural network, a module combining a residual network and a channel attention mechanism is innovatively adopted, which has stronger selectivity for input features, reduces the error of the model, and makes the convergence speed of the model faster, thereby obtaining better sound source localization accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound source localization, and particularly to a sound source localization method based on a four-microphone array and deep learning. Background Art

[0002] If one is in a noisy environment for a long time, it is very harmful to human health. Currently, the control of noise mainly starts from three aspects: the noise source, the noise propagation path, and the protection of the recipient. The most direct and effective method is to control noise at the source of noise generation. No matter which noise control method is adopted, the first thing to do is to figure out the sound source position of the main noise source, and then take corresponding inspection and control measures. Among them, the non-contact and long-distance microphone array technology has become the focus of research and has been widely used because it can directly visually identify and locate the noise source.

[0003] In recent years, with the rapid development of artificial intelligence technology, the sound source localization algorithm based on deep learning has become a research hotspot. Currently, the most popular is the sound source localization method based on the convolutional recurrent neural network, which is often used for localization in complex acoustic environments. This type of method models various acoustic signal feature parameters and constructs a mapping relationship between the sound source position and the signal feature parameters to achieve sound source localization. However, currently, the generalization ability of this type of algorithm to unknown environments (noise and reverberation) is relatively low, and its performance still needs to be further improved. Summary of the Invention

[0004] The main purpose of the present invention is to solve the problems in the prior art, such as relatively low generalization ability to unknown environments (noise and reverberation), poor screening of input features, and lack of certain robustness, and thus propose a sound source localization method based on a four-microphone array and deep learning. A sound source localization method based on a four-microphone array and deep learning provided by the present invention includes the following steps:

[0005] S1. Set up a microphone array, where the microphone array includes four microphones with a tetrahedral topological structure, and collect the sound source signal through the four-microphone array sensor to obtain the original sound field signal of the sound source point;

[0006] S2. Perform short-time Fourier transform processing on the original sound source data to convert it into STFT phase features, and input the STFT phase features into the sound source localization neural network module for training. After tuning, a trained sound source localization model is obtained;

[0007] S3. Input the original sound source data through short-time Fourier phase transform into the trained neural network model to obtain the angular information of the sound source point.

[0008] The beneficial effects provided by the present invention are:

[0009] The present invention innovatively makes a substantial improvement to the traditional convolutional recurrent neural network, adding residual blocks and an attention mechanism. We use residual blocks to replace ordinary two-dimensional convolutional layers to extract deeper features, which prevents the problems of gradient vanishing and explosion. At the same time, the attention mechanism is introduced to improve the efficiency of feature utilization. Taking the phase component after short-time Fourier transform as the input of the neural network, the phase features are used to learn the regression task of the sound source point.

[0010] The sound source localization model of the present invention adopts a main feature extraction module with residual blocks and an attention mechanism. During the model inference process, due to the simple training parameters and structure in this network block, it can better save hardware resources, facilitate hardware acceleration, and help the model to be better deployed on hardware. At the same time, connecting the recurrent layer through residual blocks and then connecting the fully connected layer helps to improve the convergence speed of the model and reduce the training error, effectively overcoming the deficiencies of the prior art. Brief Description of the Drawings

[0011] Figure 1 It is a schematic flow chart of the method of the present invention. Detailed Embodiments

[0012] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0013] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the simple structure of the system of the present invention;

[0014] A sound source localization method based on a four-microphone array and deep learning includes the following steps:

[0015] S1. Set up a microphone array, where the microphone array includes four microphones in a tetrahedral topological structure. The sound source signal is collected through the four-microphone array sensor to obtain the original sound field signal of the sound source point.

[0016] S2. Perform short-time Fourier transform processing on the original sound source data to convert it into STFT phase features, and input the STFT phase features into the sound source localization neural network module for training. After tuning, a trained sound source localization model is obtained.

[0017] S3. Input the original sound source data through short-time Fourier phase transform into the trained neural network model to obtain the angular information of the sound source point.

[0018] To facilitate the training of deep learning models, the present invention first uses the short-time Fourier transform to convert the original sound source signal collected by the microphone array sensor into a phase spectrum. Specifically, four microphone arrays are in a tetrahedral topological structure in space, with a sampling frequency of 48 kHz. According to formula (1), the original sound source signal x can be converted into a time-frequency spectrum y through the short-time Fourier transform.

[0019]

[0020] In the formula: S represents the number of sound sources, L i (b) represents the length of the Hann window, P represents the jump size between adjacent windows, and L(b) represents the width of the Hann window.

[0021] The reason for converting the original audio signal into a time-frequency image is that the result of STFT contains rich phase information. Therefore, the sound source position neural network module can capture the phase transformation between different channels and thus obtain accurate sound source position information.

[0022] The neural network module includes a two-dimensional convolutional block, a residual block, an attention block, a recurrent block, and a fully connected block.

[0023] The processing process of the sound source localization neural network module is as follows:

[0024] The STFT phase features pass through the two-dimensional convolutional block to obtain the input feature m; the input feature m passes through the residual block to obtain the superimposed information N. Among them, after the residual block processes the input feature m using formula (2) and then adds it to the input feature m, formula (2) is as follows:

[0025] N = F(m, ω) + m (2)

[0026] Among them, ω represents the weight;

[0027] The attention block is used to select the time-frequency channels of the superimposed information N to amplify the useful time-frequency information, as shown in formula (3):

[0028] O = αSigmoid{Conv[Pooling(N)]} (3)

[0029] Among them, α represents the correction coefficient, Sigmoid represents the Sigmoid function, Conv represents convolution, and Pooling represents global average pooling;

[0030] The deeper the network, the more difficult the training is because small transformations of network parameters will amplify the output and increase the cost of errors (i.e., loss). Network depth is crucial in challenging tasks. Deeper models not only perform well in classification tasks but are also very important for regression. The deeper the network, the easier the task becomes. The sound source localization model introducing the residual network can effectively solve the problem between the number of network layers and gradient vanishing or explosion.

[0031] Meanwhile, adding an attention mechanism enhances the learning ability of the model, improves the convergence speed of the model, and reduces the training error.

[0032] Specifically, the sound source localization neural network module is trained by the BP training method, and the mean squared error (MSE) is used to calculate the difference between the output sound source position and the actual sound source position to facilitate the optimization of the predicted value of the output. Here, k represents the number of samples, y t represents the actual sound source position, and y p represents the predicted sound source position.

[0033] The neural network parameters are continuously adjusted according to the cost function to iterate the deep learning model to find the optimal model.

[0034]

[0035] Finally, the collected sound signal is converted into a short-time Fourier transform phase spectrum and input into the trained optimal model to obtain the sound source position. Additionally, other tasks such as fault troubleshooting or detection can be performed in combination with the final sound source position.

[0036] The beneficial effects of the present invention are as follows:

[0037] The present invention innovatively makes significant improvements to the traditional convolutional recurrent neural network by adding residual blocks and an attention mechanism. We use residual blocks to replace ordinary two-dimensional convolutional layers to extract deeper features, which prevents gradient vanishing and explosion problems. At the same time, the attention mechanism is introduced to improve the feature utilization efficiency. Using the phase component after short-time Fourier transform as the input of the neural network, the phase features are used to learn the regression task of the sound source point.

[0038] The sound source localization model of the present invention adopts a main feature extraction module with residual blocks plus an attention mechanism. During the model inference process, due to the simple training parameters and structure in this network block, it can better save hardware resources, facilitate hardware acceleration, and help the model to be better deployed on hardware. At the same time, connecting the recurrent layer and then the fully connected layer through residual blocks helps to improve the convergence speed of the model and reduce the training error, effectively overcoming the deficiencies of the prior art.

[0039] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A sound source localization method based on a four-microphone array and deep learning, characterized in that: It includes the following steps: S1. Set up a microphone array, which includes four microphones in a tetrahedral topological structure. The sound source signal is collected through the four-microphone array sensor to obtain the original sound field signal of the sound source point; S2. Perform short-time Fourier transform processing on the original sound source data to convert it into STFT phase features, and input the STFT phase features into the sound source localization neural network module for training. After tuning, a trained sound source localization model is obtained; S3. Input the original sound source data through short-time Fourier phase transform into the trained neural network module to obtain the angle information of the sound source point; The backbone network of the sound source localization neural network module is the Res-eca network, including: two-dimensional convolution block, residual block, attention block, cyclic block and fully connected block; The processing process of the sound source localization neural network module is as follows: The STFT phase features pass through the two-dimensional convolution block to obtain the input feature m; the input feature m passes through the residual block to obtain the superimposed information N. Among them, the residual block processes the input feature m using formula (2) and then adds it to the input feature m. Formula (2) is as follows: N = F(m, ω) + m (2) where ω represents the weight; Use the attention block to perform time-frequency channel selection on the superimposed information N to amplify the useful time-frequency information, as shown in formula (3): O = αSigmoid{Conv[Pooling(N)]} where α represents the correction coefficient, Sigmoid represents the Sigmoid function, Conv represents convolution, and Pooling represents global average pooling; The useful time-frequency information passes through the cyclic block and the fully connected block to obtain the output prediction value.

2. A sound source localization method based on a four-microphone array and deep learning according to claim 1, characterized in that: The process of performing short-time Fourier transform processing in step S2 is as follows: The original sound source signal x is converted into a time-frequency image y through short-time Fourier transform according to formula (1): Where: S represents the number of sound sources, x i (b) represents the length of the Hann window, P represents the jump size between adjacent windows, and L(b) represents the width of the Hann window.

3. A sound source localization method based on a four-microphone array and deep learning according to claim 1, characterized in that: The specific process of obtaining the trained sound source localization model in step S2 is as follows: In step S2, the sound source localization neural network module is trained by the backpropagation training method of the neural network. First, calculate the difference between the output sound source position and the actual sound source position. According to this difference and each gradient, adjust the training parameters. Finally, continuously update each parameter according to the cost function through loop iteration to minimize the difference, and finally obtain the trained sound source localization model.

4. A sound source localization method based on a four-microphone array and deep learning according to claim 3, characterized in that: The mean squared error (MSE) is used to calculate the difference between the output sound source position and the actual sound source position. The formula is as follows: where k represents the number of samples, y t represents the actual sound source position, and y i represents the predicted sound source position.