Sound Imaging Method of Spiral Microphone Array Based on Multi-Task Deep Learning Network
By using a method of combining multi-task deep learning networks with spiral microphone arrays in acoustic imaging technology, the accuracy and anti-interference problems of traditional sound source positioning technology in complex acoustic environments are solved, and high-resolution and real-time sound source positioning and imaging are achieved.
Patent Information
- Application Number
- CN202411672346.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Traditional sound source positioning and imaging technology has problems such as low positioning accuracy, slow processing speed and weak anti-interference ability in complex acoustic environments with high noise, and deep learning-based methods are still insufficient in visual processing of sound pressure distribution.
The spiral microphone array based on multi-task deep learning network is adopted to extract the real and imaginary features of the spectrum graph through short-time Fourier transform (STFT), and combined with the dual attention network module, local features and global dependencies are integrated, deep-level features are further extracted through the convolutional layer and the bidirectional gated cyclic unit layer, and finally the sound source position and sound pressure distribution information are extracted through the full connection layer, and the visual results are integrated with the images captured by the camera.
Real-time sound source positioning and imaging under fewer array elements is realized, with high spatial resolution and strong anti-interference ability, and can work effectively in a high noise environment.
Smart Images

Figure CN119165446B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of acoustic imaging, and particularly relates to an acoustic imaging method for a spiral microphone array based on a multi-task deep learning network. Background Art
[0002] Traditional sound source localization and imaging technologies mainly rely on beamforming algorithms. Although such methods can achieve basic sound source localization functions to a certain extent, they generally have a high dependence on the number of array elements and the shape of the array distribution, resulting in problems such as low localization accuracy and slow processing speed. Especially in complex acoustic environments with high noise, the anti-interference ability of these technologies is relatively weak. In recent years, with the rapid development of deep learning technologies, research on sound source localization and imaging using neural network models has gradually increased. Such methods can overcome the limitations of traditional algorithms to a certain extent. However, most current deep learning-based sound source localization technologies mainly focus on the extraction of signal position information and are still insufficient in the visualization processing of sound pressure distribution, so they cannot comprehensively present the sound field information. Summary of the Invention
[0003] To solve the above problems, the present invention discloses an acoustic imaging method for a spiral microphone array based on a multi-task deep learning network, which realizes real-time sound source localization and imaging under the condition of fewer array elements, and has high spatial resolution and strong anti-interference ability.
[0004] To achieve the above object, the technical solution of the present invention is as follows:
[0005] An acoustic imaging method for a spiral microphone array based on a multi-task deep learning network, comprising the following steps:
[0006] (1) Using the short-time Fourier transform (STFT) to convert the audio signal collected by the spiral microphone array into a spectrogram, and respectively extracting the real part component and the imaginary part component of the spectrogram as feature inputs;
[0007] (2) Respectively inputting the real part component and the imaginary part component into their respective dual attention network modules, and adaptively fusing local features and global dependencies through residual blocks, multi-spectrum attention modules, and frame attention modules;
[0008] (3) Concatenating and fusing the output feature maps of the dual attention network modules of the real part stream and the imaginary part stream, and further extracting deep features through subsequent convolutional layers and bidirectional gated recurrent unit layers;
[0009] (4) Extracting sound source position and sound pressure distribution information through two fully connected layers respectively;
[0010] (5) Finally, integrating the sound pressure distribution information with the image captured by the camera to generate a visualization result.
[0011] Further, the audio signal STFT processing method in step (1) is as follows:
[0012] First, from the audio signals obtained from 12 microphones, spectrograms are extracted using a short-time Fourier transform (STFT) of 2F points. During the extraction process, a Hamming window of length 2F is adopted, and an overlap rate of 50% is set. The extracted imaginary and magnitude components are respectively used as the feature inputs for the corresponding dual attention network modules in the dual-stream structure, and their dimensions are represented as T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones.
[0013] Further, the modeling method of the dual attention network module in step (2) is as follows:
[0014] The dual attention network module inputs the feature maps processed by the residual block into the multi-spectrum attention module and the frame attention module respectively. The multi-spectrum attention module is used to model the interactions between different frequency points and the correlations between the signals captured by different microphones; the frame attention module is used to capture the interdependencies between different time points. Finally, the outputs of the two attention modules are fused together by element-wise summation.
[0015] In the dual attention network module, the structure of the residual block is as follows:
[0016] First, the number of channels of the input feature map is reduced by the first 1x1 convolutional layer; then, feature extraction is performed by the 3x3 convolutional layer; next, the number of channels of the feature map is restored by the second 1x1 convolutional layer; finally, the processed feature map and the original input feature map are fused together by element-wise summation to obtain the final output.
[0017] In the dual attention network module, the structure of the multi-spectrum attention module is as follows:
[0018] The size of the input feature map A is T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, the feature map A generates two new feature maps B and C through a convolutional layer, and the sizes of these two new feature maps are also T×F×M. Then, the feature map B is reshaped into a two-dimensional matrix of T×(F*M), and the feature map C is reshaped into a two-dimensional matrix of (F*M)×T. After that, matrix multiplication is performed on B and C, and the softmax function is applied to calculate the attention matrix X, with a size of (F*M)×(F*M). Additionally, the feature map A generates a new feature map D through another convolutional layer, with a size of T×F×M, and then it is reshaped into T×(F*M). After that, matrix multiplication is performed on the feature map D and the attention matrix X, and the result is reshaped back into the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β, and element-wise addition is performed with the original feature map A to obtain the final output feature map E.
[0019] In the dual attention network module, the structure of the frame attention module is as follows:
[0020] The size of the input feature map A is T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, the feature map A directly generates three new feature maps B, C, and D, and the sizes of these three new feature maps are also T×F×M. Then, the feature map B is reshaped into a two-dimensional matrix of (F*M)×T, and the feature map C is reshaped into a two-dimensional matrix of T×(F*M). After that, matrix multiplication is performed on B and C, and the softmax function is applied to calculate the attention matrix X, with a size of T×T. Then, matrix multiplication is performed on the attention matrix X and the feature map D with a size reshaped into T×(F*M), and the result is reshaped back into the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β, and element-wise addition is performed with the original feature map A to obtain the final output feature map E.
[0021] Furthermore, the feature fusion and extraction method in step (3) is as follows:
[0022] The output feature maps of the dual attention network modules of the real part stream and the imaginary part stream have a size of T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. The two are concatenated and fused, and deeper features are further extracted through the subsequent three convolutional layers and two bidirectional gated recurrent unit layers.
[0023] In the feature fusion and extraction method, the structure of the convolutional module is as follows:
[0024] This convolutional module has a total of three convolutional layers. Each convolutional layer contains P 3×3 filters that operate on the time-frequency-channel axis. Subsequently, the Batch Normalization technique is adopted to standardize the output activation values, and the Rectified Linear Unit (ReLU) is used as the activation function to capture the local translation-invariant features in the spectrogram. Finally, a max-pooling operation along the frequency axis is employed to reduce the dimension while keeping the time series length T unchanged. After three convolutional layers, the output dimension becomes T×2×P, where 2 represents the frequency dimension after max-pooling, and P is the number of filters.
[0025] In the feature fusion and extraction method, the structure of the bidirectional gated recurrent unit module is as follows:
[0026] This model adopts a design of two layers of bidirectional gated recurrent units (Bi-GRUs) to capture complex temporal dependencies. Each layer of Bi-GRU has Q nodes, and the hyperbolic tangent (tanh) activation function is used internally to update the hidden state, and the information flow is controlled through the gating mechanism (update gate and reset gate). After being processed by two layers of Bi-GRUs, the output dimension becomes T×Q.
[0027] Furthermore, the method for extracting the sound source position information and sound pressure distribution information in step (4) is as follows:
[0028] This method extracts the sound source position information and sound pressure distribution information through two fully connected layers respectively. Specifically, for the extraction of the sound source position information, two fully connected layers are used. The first layer has B nodes and the activation function is Linear, and the second layer has 3 nodes, corresponding to the position coordinates on the x, y, and z axes respectively, and the output values are mapped to between -1 and 1 through the Tanh activation function to standardize the position coordinates. For the extraction of the sound pressure distribution information, a similar two-layer structure is also adopted. The first layer also has B nodes and a linear activation function, while the second layer has N nodes, corresponding to the number of acoustic imaging grid points, and the ReLU activation function is used to represent the sound pressure magnitude.
[0029] Furthermore, the loss function of the network is as follows:
[0030] The loss function for the sound source coordinates is:
[0031] ,
[0032] where respectively represent the true coordinates of the sound source in three-dimensional space, respectively represent the predicted values of the coordinates of the sound source in three-dimensional space after passing through the model;
[0033] The loss function for the sound pressure distribution is:
[0034] ,
[0035] where N is the number of acoustic imaging grid points, is the true value of the sound pressure level at the i-th grid point, is the predicted value of the sound pressure level at the i-th grid point;
[0036] The total loss function is
[0037] ,
[0038] is the weight of the loss function of the sound source coordinates in the total loss function, is the weight of the loss function of the sound pressure distribution in the total loss function.
[0039] Furthermore, the acoustic imaging method in step (5) is as follows:
[0040] First, ensure the time synchronization of the sound pressure data and the image data. Then, according to the relative position and attitude of the microphone array and the camera in the audiovisual acquisition system, align the sound pressure heat map with the image space. Finally, overlay the sound pressure heat map on the image and adjust the transparency to ensure that the sound pressure information is clearly visible and does not obscure the background details, realizing the real-time visualization of the sound pressure distribution of the sound source;
[0041] where the structure of the audiovisual acquisition system is a 3-arm logarithmic spiral microphone array, and the camera is placed at the center of the microphone array;
[0042] The logarithmic spiral is represented in polar coordinates:
[0043] ,
[0044] where is the polar radius, is the polar angle, is a constant coefficient;
[0045] The arc length of the spiral between adjacent microphones increases in a geometric progression, and the ratio is ; for the logarithmic spiral with the maximum polar radius of , when , the microphones on the spiral are evenly distributed; when , the polar coordinates of the
[0046] -th microphone are:
[0047] ,
[0048] where is the The polar radius of a microphone in polar coordinates, is the number of microphones on a single cantilever. For a spiral array with 3 rotating arms, the array panel is divided into 3 equal parts by radians, and the entire array coordinates are obtained by rotating and replicating the microphone coordinate points.
[0049] The beneficial effects of the present invention are:
[0050] A sound imaging method for a spiral microphone array based on a multi-task deep learning network according to the present invention realizes real-time sound source localization and imaging under the condition of fewer array elements by introducing a two-stream structure with the real and imaginary components of the spectrogram as feature inputs and respectively combining dual attention network modules, and has high spatial resolution and strong anti-interference ability. Description of the Drawings
[0051] Figure 1 is the implementation flowchart of the method of the present invention;
[0052] Figure 2 is the structural diagram of the residual block in the method of the present invention;
[0053] Figure 3 is the structural diagram of the multi-spectrum attention module in the method of the present invention;
[0054] Figure 4 is the structural diagram of the frame attention module in the method of the present invention;
[0055] Figure 5 is the schematic diagram of the acquisition system in the method of the present invention. Detailed Embodiments
[0056] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0057] As Figure 1 shown, a sound imaging method for a spiral microphone array based on a multi-task deep learning network according to the present invention includes the following steps:
[0058] (1) Using the short-time Fourier transform (STFT) to convert the audio signal collected by the spiral microphone array into a spectrogram, and respectively extracting the real component and the imaginary component of the spectrogram as feature inputs;
[0059] (2) Respectively inputting the real component and the imaginary component into their respective dual attention network modules, and adaptively fusing local features and global dependency relationships through residual blocks, multi-spectrum attention modules and frame attention modules;
[0060] (3) Concatenate and fuse the output feature maps of the dual attention network modules for the real part stream and the imaginary part stream, and further extract deep features through subsequent convolutional layers and bidirectional gated recurrent unit layers;
[0061] (4) Extract the sound source position and sound pressure distribution information through two fully connected layers respectively;
[0062] (5) Finally, integrate the sound pressure distribution information with the images captured by the camera through the imaging module to generate a visualization result.
[0063] Further, the audio signal STFT processing method in step (1) is as follows:
[0064] First, from the audio signals obtained from 12 microphones, use the short-time Fourier transform (STFT) of 2F points to extract the spectrogram. During the extraction process, a Hamming window of length 2F is adopted, and an overlap rate of 50% is set. The extracted imaginary part and amplitude components are respectively used as the feature inputs for the corresponding dual attention network modules in the two-stream structure, and their dimensions are represented as T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones.
[0065] The modeling method of the dual attention network module in step (2) is as follows:
[0066] The dual attention network module inputs the feature maps processed by the residual block into the multi-spectrum attention module and the frame attention module respectively. The multi-spectrum attention module is used to model the interaction between different frequency points and the correlation between the signals captured by different microphones; the frame attention module is used to capture the interdependence between different time points. Finally, the outputs of the two attention modules are fused together by element-wise summation.
[0067] As Figure 2 shown, in the dual attention network module, the structure of the residual block is as follows:
[0068] First, reduce the number of channels of the input feature map through the first 1x1 convolutional layer; then, perform feature extraction through the 3x3 convolutional layer; next, restore the number of channels of the feature map through the second 1x1 convolutional layer; finally, fuse the processed feature map with the original input feature map by element-wise summation to obtain the final output.
[0069] As Figure 3 shown, in the dual attention network module, the structure of the multi-spectrum attention module is as follows:
[0070] The size of the input feature map A is T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, the feature map A generates two new feature maps B and C through a convolutional layer, and the sizes of these two new feature maps are also T×F×M. Then, the feature map B is reshaped into a two-dimensional matrix of T×(F*M), and the feature map C is reshaped into a two-dimensional matrix of (F*M)×T. After that, matrix multiplication is performed on B and C, and the softmax function is applied to calculate the attention matrix X, with a size of (F*M)×(F*M). Additionally, the feature map A generates a new feature map D through another convolutional layer, with a size of T×F×M, and then it is reshaped into T×(F*M). After that, matrix multiplication is performed on the feature map D and the attention matrix X, and the result is reshaped back into the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β, and element-wise addition is performed with the original feature map A to obtain the final output feature map E.
[0071] As Figure 4 shown, in the dual attention network module, the structure of the frame attention module is as follows:
[0072] The size of the input feature map A is T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, the feature map A directly generates three new feature maps B, C, and D, and the sizes of these three new feature maps are also T×F×M. Then, the feature map B is reshaped into a two-dimensional matrix of (F*M)×T, and the feature map C is reshaped into a two-dimensional matrix of T×(F*M). After that, matrix multiplication is performed on B and C, and the softmax function is applied to calculate the attention matrix X, with a size of T×T. Then, matrix multiplication is performed on the attention matrix X and the feature map D whose size is reshaped into T×(F*M), and the result is reshaped back into the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β, and element-wise addition is performed with the original feature map A to obtain the final output feature map E.
[0073] The feature fusion and extraction method in step (3) is as follows:
[0074] The output feature maps of the dual attention network modules of the real part stream and the imaginary part stream have a size of T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. The two are concatenated and fused, and deep features are further extracted through the subsequent three convolutional layers and two bidirectional gated recurrent unit layers.
[0075] In the feature fusion and extraction method, the structure of the convolutional module is as follows:
[0076] This convolutional module has a total of three convolutional layers. Each convolutional layer contains P 3×3 filters that operate on the time-frequency-channel axis. Subsequently, the Batch Normalization technique is adopted to standardize the output activation values, and the Rectified Linear Unit (ReLU) is used as the activation function to capture local translation-invariant features in the spectrogram. Finally, a max-pooling operation along the frequency axis is employed to reduce the dimension while keeping the time series length T unchanged. After passing through the three convolutional layers, the output dimension becomes T×2×P, where 2 represents the frequency dimension after max-pooling and P is the number of filters.
[0077] In the feature fusion and extraction method, the structure of the bidirectional gated recurrent unit module is as follows:
[0078] This model adopts a design of two layers of bidirectional gated recurrent units (Bi-GRUs) to capture complex temporal dependencies. There are Q nodes in each layer of Bi-GRU, and the hyperbolic tangent (tanh) activation function is used internally to update the hidden state, and the information flow is controlled through the gating mechanism (update gate and reset gate). After being processed by two layers of Bi-GRUs, the output dimension becomes T×Q.
[0079] The method for extracting the sound source position information and sound pressure distribution information in step (4) is as follows:
[0080] This method extracts the sound source position information and sound pressure distribution information through two fully connected layers respectively. Specifically, for the extraction of the sound source position information, two fully connected layers are used. The first layer has B nodes and the activation function is Linear, and the second layer has 3 nodes, corresponding to the position coordinates on the x, y, and z axes respectively, and the output values are mapped to between -1 and 1 through the Tanh activation function to standardize the position coordinates. For the extraction of the sound pressure distribution information, a similar two-layer structure is also adopted. The first layer also has B nodes and a linear activation function, while the second layer has N nodes, corresponding to the number of acoustic imaging grid points, and the ReLU activation function is used to represent the sound pressure magnitude.
[0081] The loss function of the network is as follows:
[0082] The loss function for the sound source coordinates is:
[0083] ,
[0084] where represent the true coordinates of the sound source in three-dimensional space respectively, represent the predicted values of the coordinates of the sound source in three-dimensional space after passing through the model respectively;
[0085] The loss function for the sound pressure distribution is:
[0086] ,
[0087] where N is the number of acoustic imaging grid points, is the true value of the sound pressure level at the i-th grid point, is the predicted value of the sound pressure level at the i-th grid point;
[0088] The total loss function is
[0089] ,
[0090] is the weight of the loss function of the sound source coordinates in the total loss function, is the weight of the loss function of the sound pressure distribution in the total loss function.
[0091] The acoustic imaging method in step (5) is as follows:
[0092] First, ensure the time synchronization of the sound pressure data and the image data, and then align the sound pressure heat map with the image space according to the relative position and attitude of the microphone array and the camera in the audio-visual acquisition system; finally, overlay the sound pressure heat map on the image and adjust the transparency to ensure that the sound pressure information is clearly visible and does not obscure the background details, realizing the real-time visualization of the sound pressure distribution of the sound source;
[0093] Among them, the structure of the audio-visual acquisition system is a 3-arm logarithmic spiral microphone array, and the camera is placed at the center of the microphone array;
[0094] The logarithmic spiral is represented in polar coordinates:
[0095] ,
[0096] where is the polar radius, is the polar angle, is the constant coefficient;
[0097] The arc length of the spiral between adjacent microphones increases in a geometric progression, and the ratio is ; for the logarithmic spiral with the maximum polar radius of , when , the microphones on the spiral are evenly distributed; when, the th microphone's polar coordinates are:
[0098] ,
[0099] ,
[0100] where is the polar radius of the th microphone in polar coordinates, is the number of microphones on a single cantilever. For a spiral array with 3 cantilevers, the array panel is equally divided into 3 parts by radians, and the entire array coordinates are obtained by rotating and copying the microphone coordinate points.
[0101] The present invention introduces a two-stream structure with the real and imaginary components of the spectrogram as feature inputs, and combines dual attention network modules respectively, achieving real-time sound source localization and imaging under the condition of fewer array elements, with high spatial resolution and strong anti-interference ability.
Claims
1. An acoustic imaging method of a spiral microphone array based on a multi-task deep learning network, characterized in that: The following steps are involved: (1) Using STFT, the audio signal collected by the spiral microphone array is converted into a spectrogram, and the real and imaginary components of the spectrogram are extracted as feature inputs. First, the spectrogram is extracted from the audio signals obtained from 12 microphones using the 2F-point STFT; During the extraction process, a Hamming window with a length of 2F is used, and an overlap rate of 50% is set; the extracted imaginary part and amplitude components are used as feature inputs of their respective dual attention network modules, and their dimensions are expressed as T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones; (2) The real and imaginary components are input into their respective dual attention network modules, and the local features and global dependencies are adaptively fused through the residual block, multi-spectral attention module and frame attention module; The dual attention network module inputs the feature map processed by the residual block to the multi-spectral attention module and the frame attention module respectively; the multi-spectral attention module is used to model the interaction between different frequency points and the correlation between the signals captured by different microphones; The frame attention module is used to capture the interdependencies between different time points. Finally, the outputs of the two attention modules are fused together by element-wise summation. The structure of the frame attention module is as follows: The input feature map A has a size of T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, the feature map A directly generates three new feature maps B, C, and D, and the sizes of these three new feature maps are also T×F×M. Then the feature map B is reshaped into a two-dimensional matrix of (F*M)×T, and the feature map C is reshaped into a two-dimensional matrix of T×(F*M). After that, matrix multiplication is performed on B and C and the softmax function is applied to calculate the attention matrix X, which has a size of T×T. After that, the attention matrix X is matrix multiplied with the feature map D reshaped to a size of T×(F*M), and the result is reshaped back to the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β and element-by-element addition is performed on the original feature map A to obtain the final output feature map E. (3) The feature maps output by the dual attention network modules of the real and imaginary streams are concatenated and fused, and the deep features are further extracted through subsequent convolutional layers and Bi-GRU layers; (4) Extract the sound source position and sound pressure distribution information respectively through two fully connected layers; (5) Finally, the sound pressure distribution information is integrated with the image captured by the camera to generate a visualization result.
2. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The structure of the residual block is as follows: First, the number of channels of the input feature map is reduced through the first 1x1 convolution layer; then, features are extracted through a 3x3 convolution layer; then, the number of channels of the feature map is restored through the second 1x1 convolution layer; finally, the processed feature map is fused with the original input feature map by element-wise summation to obtain the final output.
3. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The structure of the multi-spectral attention module is as follows: The input feature map A has a size of T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. First, feature map A generates two new feature maps B and C through a convolutional layer, and the size of these two new feature maps is also T×F×M. Then feature map B is reshaped into a two-dimensional matrix of T×(F*M), and feature map C is reshaped into a two-dimensional matrix of (F*M)×T. After that, matrix multiplication is performed on B and C and the softmax function is applied to calculate the attention matrix X, which has a size of (F*M)×(F*M). In addition, feature map A generates a new feature map D through another convolutional layer, which has a size of T×F×M and is then reshaped to T×(F*M). After that, matrix multiplication is performed on feature map D and attention matrix X, and the result is reshaped back to the shape of T×F×M. Subsequently, it is multiplied by a scale parameter β and element-wise added to the original feature map A to obtain the final output feature map E.
4. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The feature fusion and extraction method in step (3) is as follows: The dual attention network module of the real stream and the imaginary stream outputs feature maps of size T×F×M, where T represents the number of time frames, F represents the number of positive frequency components, and M represents the number of microphones. The two are concatenated and fused, and the subsequent convolution module and bidirectional gated recurrent unit module are used to further extract deep features. The convolution module has three convolutional layers; each convolutional layer contains P 3×3 filters, which work on the time-frequency-channel axis; batch normalization technology is then used to standardize the output activation value, and the rectified linear unit ReLU is used as the activation function to capture the local translation invariant features in the spectrogram; finally, the maximum pooling operation along the frequency axis is used to reduce the dimension while keeping the time series length T unchanged; after three convolutional layers, the output dimension becomes T×2×P, where 2 represents the frequency dimension after maximum pooling, P is the number of filters, The bidirectional gated recurrent unit module adopts a two-layer Bi-GRU design to capture complex temporal dependencies. There are Q nodes in each layer of Bi-GRU, which uses a hyperbolic tangent activation function to update the hidden state and controls the flow of information through a gating mechanism. After two layers of Bi-GRU processing, the output dimension becomes T×Q.
5. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The method for extracting the sound source position information and the sound pressure distribution information in step (4) is as follows: This method extracts the location information and sound pressure distribution information of the sound source through two fully connected layers respectively; specifically, for the extraction of the sound source location information, two fully connected layers are used, where the first layer has B nodes and the activation function is Linear, and the second layer has 3 nodes, corresponding to the position coordinates of the x, y, and z axes respectively, and the output value is mapped to between -1 and 1 through the Tanh activation function, thereby standardizing the position coordinates; for the extraction of sound pressure distribution information, a similar two-layer structure is also used, the first layer also has B nodes and a linear activation function, and the second layer has N nodes, corresponding to the number of acoustic imaging grid points, and the ReLU activation function is used to represent the sound pressure.
6. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The loss function of the network is as follows: The loss function of the sound source coordinates is: Among them, x, y, and z represent the real coordinates of the sound source in three-dimensional space. They respectively represent the predicted coordinate values of the sound source passing through the model in three-dimensional space; The loss function of the sound pressure distribution is: Where N is the number of acoustic imaging grid points, p i is the true value of the sound pressure level at the ith grid point, is the predicted value of the sound pressure level at the i-th grid point; The total loss function is L=w p L p +w spl L spl w p is the weight of the loss function of the sound source coordinates in the total loss function, w spl is the weight of the loss function of sound pressure distribution in the total loss function.
7. The acoustic imaging method of a spiral microphone array based on a multi-task deep learning network according to claim 1, characterized in that: The acoustic imaging method in step (5) is as follows: First, ensure the time synchronization of sound pressure data and image data. Then, align the sound pressure heat map with the image space according to the relative position and posture of the microphone array and the camera in the audio-visual acquisition system. Finally, overlay the sound pressure heat map on the image and adjust the transparency to ensure that the sound pressure information is clearly visible without blocking the background details, so as to achieve real-time visualization of the sound pressure distribution of the sound source. The structure of the audio-visual acquisition system is a 3-arm logarithmic spiral microphone array, and the camera is placed at the center of the microphone array; The logarithmic spiral is represented by polar coordinates: r=a·e kθ Where r is the polar diameter, θ is the polar angle, and a and k are constant coefficients; The arc length of the spiral between two adjacent microphones increases in geometric proportion, and the ratio is q. For a logarithmic spiral with a maximum polar diameter of R, when q = 1, the microphones on the spiral are evenly distributed. When q ≠ 1, the polar coordinates of the mth microphone are: Among them, r m is the polar diameter of the mth microphone in polar coordinates, M is the number of microphones on a single cantilever, and for a spiral array with three spiral arms, the array panel is divided into three equal parts of 2π / 3 radians, and the entire array coordinates are obtained by rotating and copying the microphone coordinate points.
Citation Information
Patent Citations
Audio-visual voice noise reduction method based on multi-mode gating lifting model
CN116013297A
Sound source localization method based on four-microphone array and deep learning
CN116106827A
Multi-sound-source positioning and imaging method of distributed microphone array
CN116400296A