A sound source identification method based on array expansion and SE-moving inverse bottleneck convolutional neural network

By combining EAG-U-Net and SE-MBCNet networks, the problem of insufficient accuracy and robustness of traditional sound source localization methods in complex environments is solved, and the effect of improving the accuracy and clarity of sound source localization is achieved without increasing the physical array.

CN119767201BActive Publication Date: 2025-11-14ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411877044.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-11-14
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Traditional sound source localization methods suffer from reduced accuracy and robustness in complex environments with high noise levels or a limited number of arrays. Furthermore, increasing the number of sensors leads to higher hardware costs and computational burdens, making it difficult to improve localization accuracy without increasing the physical array.

Method used

A sound source identification method based on array extension and SE moving inverse bottleneck convolutional neural network is adopted. The sound source distribution map predicted by the 18-array MUSIC algorithm is converted into the sound source distribution map predicted by the 64-array MUSIC algorithm through the EAG-U-Net data conversion model. The image processing is combined with the SE-MBCNet network to improve the localization accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of sound source localization, reduces the number of microphone arrays required, enhances its application capabilities in complex acoustic environments, and improves the clarity and accuracy of sound source distribution maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119767201B_ABST
    Figure CN119767201B_ABST
Patent Text Reader

Abstract

This invention relates to a sound source identification method based on array expansion and SE (Search Engine Attention) mobile inverse bottleneck convolutional neural network, comprising: acquiring an 18-array sound source distribution map calculated using the MUSIC algorithm based on the sound pressure cross-spectrum matrix; inputting the 18-array sound source distribution map into an EAG-U-Net data conversion model to convert the 18-array sound source distribution map into a 64-array sound source distribution map, achieving the effect of expanding the microphone array. This model introduces the EAG mechanism to optimize model feature selection in the spatial and channel dimensions, improving the model's representational ability and accuracy in the data conversion process, and generating an acoustic imaging map after data conversion; inputting the 64-array sound source distribution map into a mobile inverse bottleneck convolutional neural network model based on the SE attention mechanism for feature extraction and image reconstruction to obtain a predicted sound source distribution map; and performing local maxima detection using the predicted sound source distribution map to obtain sound source localization and intensity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of sound source localization and deep learning technology, and proposes a sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network. Background Technology

[0002] Sound source localization technology has important applications in mechanical fault diagnosis, environmental noise control, and automotive NVH development. In industrial environments, equipment fault sounds are often masked by other noises, making it difficult to detect problems in a timely manner. Sound source localization can accurately identify weak fault signals and effectively locate the noise source, providing support for noise control and equipment maintenance. With the development of microphone array and beamforming technologies, the accuracy and practicality of sound source localization are constantly improving. Meanwhile, advancements in artificial intelligence and deep learning have enabled sound signal-based localization and recognition technologies to show broad prospects in areas such as speech recognition, autonomous driving, and security monitoring. However, traditional methods still face challenges in accuracy, robustness, and real-time performance, urgently requiring new technological innovations to address these issues.

[0003] Sound source localization technology has made significant progress in recent years, especially in array-based direction estimation methods. The MUSIC (Multiple Signal Classification) algorithm is widely used due to its high-precision localization capabilities. This algorithm obtains spatial information of the sound source by calculating the sound pressure cross-spectrum matrix and utilizing an array of sensors. However, in complex environments with high noise levels or a limited number of arrays, the accuracy and robustness of the MUSIC algorithm are affected. In the MUSIC algorithm, the size of the sound source distribution map typically depends on the number of microphones in the array, and training a deep learning model with more microphone arrays may lead to a decrease in localization accuracy with fewer microphone arrays. Traditional methods improve accuracy by increasing the number of array sensors, but this faces challenges such as hardware cost, array layout complexity, and computational burden. Therefore, how to extend the array's receiving range through algorithms without increasing the physical array has become an important research direction in the field of sound source localization.

[0004] With the rapid development of deep learning technology, its application in signal processing has gradually shown great potential. In sound source identification, deep learning-based methods are mainly divided into two categories: grid-based and gridless. Grid-based methods divide the sound source plane into multiple grids and identify the sound source within each grid. For example, Ma et al. applied convolutional neural networks (CNNs) to sound pressure distribution maps, combining deep learning and beamforming algorithms for sound source identification for the first time; Xu et al. used densely connected convolutional neural networks (DenseNet) to generate sound pressure distribution maps from cross-spectral matrices (CSMs), achieving high-resolution sound source identification and capable of identifying up to 25 sound sources at a single frequency. To further improve accuracy, Lee et al. proposed a new objective function dependent on the relative positional relationship between grid points and sound sources, and combined it with a fully convolutional neural network (FCN) with an encoder-decoder structure, achieving high-resolution and multi-source identification. The combination of these methods not only improves the accuracy of sound source localization but also enhances the robustness of the system in complex environments. Deep convolutional neural networks based on the U-Net structure further improve the accuracy of sound source localization. Furthermore, the combination of self-attention and gated attention mechanisms has been applied to image enhancement and signal processing tasks, significantly improving performance.

[0005] The size of the sound source distribution map, used as input for deep learning, is determined by the number of grid points. When the number of microphones decreases, the sound source features in the distribution map become less prominent. Therefore, when training a deep learning model using data generated by a larger microphone array to predict sound source distribution maps with a smaller array, the localization accuracy may decrease. To overcome this problem and enhance the versatility of deep learning sound source recognition methods, this invention proposes a sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural networks. Specifically, the EAG-U-Net data transformation model converts the sound source distribution map predicted by the 18-array MUSIC algorithm into one predicted by the 64-array MUSIC algorithm, thus solving the problem of the MUSIC algorithm's sensitivity to array geometry and the number of elements. The SE-MBCNet network combines the SE attention mechanism with a convolutional neural network, enabling more accurate extraction of key information from images, especially significantly improving the accuracy of sound source localization when processing complex sound source distribution maps. Summary of the Invention

[0006] This invention proposes a sound source recognition method based on array expansion and SE moving inverse convolutional neural network. EAG-U-Net, as a novel data conversion network, inherits the basic framework of the U-Net network (Efficient Attention Gate U-Net) and adopts an encoder-decoder structure. It effectively extracts and recovers spatial features of the image through downsampling and upsampling operations. A unique skip connection mechanism passes the feature map from the encoder to the decoder layer, improving detail preservation and sound source localization accuracy. The EAG-U-Net data conversion model transforms the sound source distribution map predicted by the 18-array MUSIC algorithm into the sound source distribution map predicted by the 64-array MUSIC algorithm, compensating for the MUSIC algorithm's dependence on array structure, microphone distribution, and microphone number to achieve array expansion. The SE-MBCNet network performs image processing on the acoustic image generated by the MUSIC algorithm, improving the robustness of the MUSIC algorithm in noisy and complex environments, and enhancing the effectiveness and performance of sound source localization.

[0007] To achieve the above objectives, the present invention provides the following solution:

[0008] A sound source recognition method based on array expansion and SE-moving inverse bottleneck convolutional neural network, the method comprising:

[0009] Step 1: Obtain the sound pressure cross-spectrum matrix, and use the sound pressure cross-spectrum matrix to calculate the 18-array sound source distribution map and the 64-array sound source distribution map of the MUSIC algorithm;

[0010] Step 2: Input the 18-array sound source distribution map into the EAG-U-Net data conversion model to obtain the converted 64-array sound source distribution map; the EAG-U-Net data conversion model is trained using the 18-array sound source distribution map as input and the 64-array sound source distribution map as labels.

[0011] Step 3: Input the transformed 64-array sound source distribution map into the SE-MBCNet network model to obtain the predicted sound source distribution map; the SE-MBCNet network model is trained using the transformed 64-array sound source distribution map as input and the target map as the label.

[0012] Step 4: Perform local maxima detection on the predicted sound source distribution map to obtain the sound source location and intensity.

[0013] Optionally, obtaining the 18-array sound source distribution map and the 64-array sound source distribution map includes:

[0014] Step 1.1: Mesh the sound source plane, discretize the sound source calculation plane to form a mesh surface, i.e., the focusing plane, and take each mesh point of the focusing plane as a potential sound source point;

[0015] Step 1.2: Arrange multiple microphone sensors in the sound field formed by the radiation of several sound sources to form a planar microphone array, i.e., a measurement plane, and make the measurement plane parallel to the sound source plane;

[0016] Step 1.3: Receive frequency sound pressure signal data collected by the sensor array, wherein the sound pressure signal data includes: sound source signal and noise signal;

[0017] Step 1.4: Preprocess the frequency sound pressure signal data, calculate the sound pressure cross spectrum matrix, and obtain the 18-array sound source distribution map and the 64-array sound source distribution map according to the MUSIC algorithm.

[0018] Optionally, obtaining the converted 64-array sound source distribution map includes:

[0019] Step 2.1: Input the 18-array sound source distribution map into the initial convolutional layer of the EAG-U-Net data conversion model. The initial convolutional layer uses two 3×3 convolutions to perform preliminary processing on the input image and extract the basic features of the 18-array sound source distribution map.

[0020] Step 2.2: Input the basic features into the encoder of the EAG-U-Net data transformation model. The encoder consists of three layers: a double convolutional layer and a pooling layer. It extracts image features continuously and performs downsampling to gradually extract the first high-level features. The double convolutional layer contains two 3×3 convolutions, a ReLU activation function, and a batch normalization layer. The pooling layer uses max pooling for downsampling to reduce the spatial dimension and increase the receptive field, thereby reducing the computational load and extracting more abstract features, namely the first high-level features.

[0021] Step 2.3: After the pooling operation in the third encoder block, an ECA attention mechanism is added. The ECA attention mechanism obtains the average value of each channel by performing global average pooling to capture global information. Based on the global information, a one-dimensional convolution operation is performed on the pooled channel information to capture the dependency relationship between channels. Based on the dependency relationship, a weight coefficient is obtained through the Sigmoid activation function. The weight coefficient is used to perform a weighted operation on each channel.

[0022] Step 2.4: Using the AG gated attention mechanism, the encoder feature map Fe and the corresponding level decoder feature map Fd are concatenated to generate a first fused feature map. The first fused feature map is then subjected to a 1×1 convolution to extract higher-level feature information, reduce the number of channels, and generate the input features of the attention map. The sigmoid activation function is used to generate an attention weight map A. The attention weight map A is multiplied element-wise with the encoder feature map Fe to apply attention weights to the encoder feature map, dynamically adjusting the weighting of the feature map. The weighted feature map is then fused with the feature map in the decoder to obtain a second fused feature map.

[0023] When restoring spatial resolution, the decoder concatenates the feature map of the encoder with the feature map of the encoder to form a first fused feature map. The first fused feature map contains the low-level features of the encoder and the high-level semantic information of the decoder.

[0024] Step 2.5: Upsample the output of the third encoder block through transposed convolution. The result and the output of the second encoder block are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the third decoder block.

[0025] The output of the third-layer decoder block is upsampled through transposed convolution. The result and the output of the first encoder block are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the second-layer decoder block.

[0026] The output of the second layer decoder block is upsampled through transposed convolution. The result and the output of the initial convolutional layer are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the first layer decoder block.

[0027] The output of the first layer decoder block is used to restore image details using a double convolutional layer to obtain a 64-array sound source distribution map after data conversion.

[0028] Optionally, obtaining the predicted sound source distribution map includes:

[0029] Step 3.1: Input the 64-array sound source distribution map after data conversion into the two-layer convolution module of the SE-MBCNet network model. The two-layer convolution module is used to perform preliminary feature extraction on the 64-array sound source distribution map after data conversion. The two-layer convolution module includes two consecutive 3×3 convolution operations, batch normalization layer BN, and activation function ReLU.

[0030] Step 3.2: Input the extracted preliminary features into the encoder module of the SE-MBCNet network model. The encoder module is used to encode multi-level semantic information and features into the feature map after preliminary feature extraction. The encoder module consists of four MBConv layers with SE attention mechanism and a max pooling layer. It extracts second-level features step by step by continuously extracting image features and downsampling them. The MBConv layer with SE attention mechanism includes: depthwise separable convolution, SE attention mechanism and residual connection.

[0031] The workflow of the depthwise separable convolution, SE attention mechanism, and residuals includes:

[0032] First, dimensionality is increased by expanding the feature dimension using 1×1 convolutions, batch normalization (BN), and the Swish activation function to increase the number of channels. Next, depthwise separable convolutions are performed, where 3×3 depthwise convolutions, combined with BN and the Swish activation function, process each channel independently to extract local features. Then, a channel attention (SE) mechanism is introduced, first using a global average pooling layer for compression to effectively capture global features and reduce computational complexity. Then, in the activation phase, weights are adaptively generated for each channel through fully connected layers and the Sigmoid activation function to enhance important features. In the scaling phase, the features of each channel are reweighted by multiplying the obtained channel weights element-wise with each channel of the original feature map. Finally, 1×1 convolutions are used to reduce the weighted feature map to the number of output channels, and residual connections are added to maintain network stability and accelerate training.

[0033] The expression formula for the MBConv layer of the SE attention mechanism is as follows:

[0034] X exp= Swish(BN(Conv1(X)))

[0035] X dw =Swish(BN(DwConv(X) exp )))

[0036] X out =X+Conv1(SE(X) dw ))

[0037] Where Swish represents self-gated activation function, BN represents batch normalization layer, Conv1 represents 1×1 convolution, DwConv represents 3×3 convolution for each channel independently, SE represents SE channel attention mechanism, and X exp X represents the result of the dimensionality upgrade operation. out The MBConv module results represent the SE attention mechanism;

[0038] Step 3.3: Input the feature map processed by the encoder module of the SE-MBCNet network model into the decoder to recover the high-resolution feature map; the decoder is divided into four stages, and the decoder block of each stage directly concatenates the feature maps with the corresponding layer of the encoder through skip connections, merging the downsampled feature map and the upsampled feature map to enhance the features of the recovered image; the encoder block consists of the MBConv layer of the SE attention mechanism and the upsampling layer, which retains detailed information, thereby avoiding the loss of important features during the upsampling process;

[0039] Step 3.4: Obtain the final predicted sound source distribution map with sound source location coordinates and intensity information through the 1×1 convolutional layer at the end of the network.

[0040] Optionally, obtaining the sound source location and intensity includes:

[0041] Step 4.1: Using a local maximum detection algorithm, identify the local maxima in each row of the data matrix of the predicted sound source distribution map and find the top three maximum values;

[0042] Step 4.2: Verify whether the maximum value is a true local maximum by comparing the target value with the values ​​in its neighborhood to determine whether the maximum value is the largest in the local range. If the maximum value is the largest in the local range, sort the detected local maxima.

[0043] Step 4.3: Based on the sorting results, use coordinate transformation technology to map the row and column indices of the maximum values ​​to the actual coordinates to calculate the location and intensity of the source.

[0044] The beneficial effects of this invention are as follows:

[0045] This invention proposes an innovative sound source localization method that combines data transformation with a deep learning model. By introducing an innovative data transformation preprocessing method, the sound source localization process is effectively optimized before model training. Specifically, EAG-U-Net data transformation is used to convert the sound source distribution map generated by the 18-array MUSIC algorithm into the prediction result of the 64-array MUSIC algorithm, thus overcoming the sensitivity of the MUSIC algorithm to changes in array geometry and the number of array elements. This method significantly reduces the number of microphone arrays required, enabling low-density arrays to simulate the accuracy of high-density arrays. Furthermore, the method performs excellently in sidelobe cancellation and main lobe width reduction, improving the clarity and accuracy of the sound source distribution map.

[0046] This invention combines deep learning technology with the SE-MBCNet network to optimize the sound source distribution map predicted by the MUSIC algorithm, effectively improving the resolution of acoustic images and the accuracy of sound source localization. In the traditional MUSIC algorithm, sound source localization is affected by noise. By introducing the deep learning SE-MBCNet network model, this invention significantly enhances the accuracy of imaging results while maintaining high-resolution sound source localization capabilities.

[0047] This invention, through training a deep learning network, can accurately predict the location and intensity of a sound source. Traditional methods rely on a large array of microphones to achieve high-precision localization, while this invention can extract sound source information using limited microphone data, significantly improving the accuracy and robustness of sound source localization and enhancing its application capabilities in complex acoustic environments. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart of a sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network according to an embodiment of the present invention;

[0050] Figure 2 This is a schematic diagram of the specific network structure of the converter EAG-U-Net according to an embodiment of the present invention; wherein, (a) is a structural diagram of the converter EAG-U-Net, (b) is a structural diagram of ECA, and (c) is a structural diagram of AG;

[0051] Figure 3 This is a schematic diagram of the specific structure of the SE-MBCNet network according to an embodiment of the present invention; wherein, (a) is the SE-MBCNet structure diagram, (b) is the MBConv structure diagram, and (c) is the SE attention mechanism structure diagram;

[0052] Figure 4 This is a schematic diagram of a sound source localization measurement model according to an embodiment of the present invention; wherein, (a) is a schematic diagram of a sound source localization measurement model, and (b) is an 18-helix channel microphone array and a 64-helix channel microphone array;

[0053] Figure 5The following are the location results of a single sound source according to an embodiment of the present invention: (a) is the location result at the single sound source position [-0.31, 0.27, 0], SNR = 20dB, f = 1000Hz, z = 1m; (b) is the location result at the single sound source position [0.17, 0.53, 0], SNR = 20dB, f = 2000Hz, z = 1m; and (c) is the location result at the single sound source position [-0.57, -0.15, 0], SNR = 20dB, f = 6000Hz, z = 1m.

[0054] Figure 6 The following are the positioning results of the dual sound source locations according to an embodiment of the present invention: (a) is the positioning result at the dual sound source locations [-0.21, -0.19, 0], [0.59, -0.19, 0], SNR = 20dB, f = 1000Hz, z = 1m; (b) is the positioning result at the dual sound source locations [0.05, -0.27, 0], [-0.25, -0.39, 0], SNR = 20dB, f = 2000Hz, z = 1m; and (c) is the positioning result at the dual sound source locations [-0.59, -0.13, 0], [0.47, -0.27, 0], SNR = 20dB, f = 6000Hz, z = 1m.

[0055] Figure 7 The diagram shows the location results of the three sound sources according to an embodiment of the present invention; (a) is the location result at the three sound source positions [-0.07, 0.53, 0], [0.35, 0.11, 0], [-0.05, -0.41, 0], SNR = 20dB, f = 1000Hz, z = 1m; (b) is the location result at the three sound source positions [-0.15, 0.19, 0], [-0.39, ... (c) shows the positioning results at the three sound source positions [0.53, 0.37, 0], [0.11, 0.17, 0], [0.45, -0.15, 0], SNR = 20dB, f = 2000Hz, z = 1m. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] like Figure 1 As shown, this embodiment discloses a sound source identification method based on array extension and SE moving inverse bottleneck convolutional neural network. The method includes: Step 1: Obtaining the sound pressure cross-spectrum matrix and calculating the 18-array sound source distribution map and the 64-array sound source distribution map of the MUSIC algorithm. Step 2: Inputting the 18-array sound source distribution map into the EAG-U-Net data conversion model to obtain the converted 64-array sound source distribution map. The EAG-U-Net data conversion model is trained using the 18-array sound source distribution map as input and the 64-array sound source distribution map as labels. Step 3: Inputting the converted 64-array sound source distribution map into the SE-MBCNet network model to obtain the predicted sound source distribution map. The SE-MBCNet network model is trained using the converted 64-array sound source distribution map as input and the target map as labels. Step 4: Performing local maxima detection on the predicted result map to obtain the sound source localization and intensity.

[0059] Specifically:

[0060] Step 1: Obtain the sound pressure cross spectrum matrix and calculate the 18-array sound source distribution map and 64-array sound source distribution map of the MUSIC algorithm.

[0061] Step 1 specifically involves:

[0062] Step 1.1: Arrange M microphone sensors within the sound field formed by the radiation from K sound sources. The M sensors form a planar microphone array, called the measurement plane, which is parallel to the sound source plane, with a distance of z between the two planes. For example... Figure 4 (a)-(b).

[0063] Step 1.2: Mesh the sound source plane. Discretize the sound source calculation plane to form a mesh surface, called the focusing plane. The focusing plane contains N mesh points, each of which is also called a focal point, and each focal point serves as a potential sound source point.

[0064] Step 1.3: Receive the frequency sound pressure signal data X collected by the sensor array, which includes the sound source signal and the noise signal, and then calculate its sound pressure cross spectrum matrix C.

[0065]

[0066] Where X = [x1, x2, ..., x M ] T H denotes the conjugate transpose of a matrix.

[0067] Step 1.4: Perform eigenvalue decomposition on the cross-spectral matrix C. Select the eigenvectors corresponding to the K largest eigenvalues ​​as the signal subspace E, and the other eigenvectors as the noise subspace.

[0068] C=UAU H

[0069] in, Let A be the eigenvector matrix. A = diag(λ1, λ2, ..., λ M Let be a diagonal eigenvalue matrix, where λ1≥λ2≥...≥λ M .

[0070]

[0071] Step 1.4: Utilizing the core steps of the MUSIC algorithm, calculate the spatial spectrum P based on the acoustic pressure cross-spectrum matrix C. Use this as input for subsequent algorithmic analysis and processing, as shown in the following formula:

[0072]

[0073] Where g is the steering vector, ||r n -r m || represents the distance from the m-th microphone to the n-th grid point, c represents the speed of sound, and f represents the frequency.

[0074] Step 1.5: Normalize the calculated spatial spectrum P and perform acoustic imaging to obtain a sound source distribution map.

[0075] Step 1.6: A target image was designed as the ground truth map for acoustic imaging. This image was calculated based on the location, intensity, and frequency of the sound source.

[0076]

[0077] Where, q k Let |r| represent the intensity of the k-th sound source. n -r k || represents the distance between the k-th sound source and the n-th grid point in the target image. η and β are two adjustable hyperparameters. η is a very small constant used to avoid zero in the denominator, and β is used to control the attenuation rate of the main lobe in the target image.

[0078] Step 2: Input the 18-array sound source distribution map into the EAG-U-Net data conversion model to obtain a 64-array sound source distribution map. The structure of the EAG-U-Net data conversion model is as follows: Figure 2 (a)-(c).

[0079] Step 2 specifically involves:

[0080] Step 2.1: Input the 18-array sound source distribution map into the initial convolutional layer of the EAG-U-Net data conversion model. The initial convolutional layer uses two 3×3 convolutions to perform preliminary processing on the input image and extract the basic features of the 18-array sound source distribution map.

[0081] Step 2.2: The basic features are input into the encoder of the EAG-U-Net data transformation model. The encoder processes the data layer by layer through three encoder blocks. Each encoder block consists of two 3×3 convolutional layers, a ReLU activation function, a batch normalization layer, and a max pooling layer. Convolutional operations extract low-level and mid-level features of the image, the ReLU activation function increases non-linearity, and the batch normalization layer accelerates training and improves stability. Max pooling reduces spatial resolution through downsampling, increases the receptive field, and extracts more abstract features. After these processes, the spatial size of the image is gradually compressed, but its high-level semantic information is effectively preserved.

[0082] Step 2.3: In the third encoder block, after the same convolution, activation, batch normalization, and pooling operations as the first two blocks, an ECA attention mechanism is added. First, global average pooling is performed to obtain the average value for each channel, used to capture global information. Then, a one-dimensional convolution operation is performed on the pooled channel information to capture the dependencies between channels. Based on these dependencies, a weight coefficient is obtained through the Sigmoid activation function. This weight coefficient is then used to weight each channel, further enhancing the network's focus on important features.

[0083] The formula for calculating the ECA attention mechanism:

[0084]

[0085] Here, Pool(·) represents global average pooling. Conv1(·) represents a one-dimensional convolution operation. σ is the sigmoid activation function, used to generate the weight coefficients for each channel.

[0086] The ECA attention mechanism is used to further enhance the feature representation at the channel level before the upsampling process of the decoder, adjust the channel weights, and obtain key image features.

[0087] Step 2.4: Using the AG gating attention mechanism, the encoder feature map Fe and the corresponding level decoder feature map Fd are concatenated to generate a first fused feature map. The first fused feature map is then subjected to a 1×1 convolution to extract higher-level feature information, reduce the number of channels, and generate the input features of the attention map. The sigmoid activation function is used to generate an attention weight map A. The attention weight map A is then multiplied element-wise with the encoder feature map Fe to apply attention weights to the encoder feature map, dynamically adjusting the weighting of the feature map. The weighted feature map is then fused with the feature map in the decoder to obtain a second fused feature map.

[0088] When restoring spatial resolution, the decoder first concatenates the feature map with the encoder's feature map to form a first fused feature map. This feature map contains the low-level features of the encoder and the high-level semantic information of the decoder.

[0089] The calculation formula for the Attention Gate (AG) module is as follows:

[0090] A=σ(Conv[F e ,F d ])

[0091] F AG =A⊙F d

[0092] Where σ is the sigmoid activation function. Conv(·) represents a 1×1 convolution. [·] represents the concatenation operation. ⊙ is the element-wise multiplication operation.

[0093] The AttentionGate module improves the feature map fusion process by dynamically adjusting the weights of each region in the input feature map, giving more attention to important regions and suppressing secondary regions, thereby improving the efficiency and accuracy of the model when processing image tasks.

[0094] Step 2.5: Upsample the output of the third encoder block through transposed convolution, and feed the result together with the output of the second encoder block into the AG gated attention mechanism for processing. The resulting feature map is then concatenated with the input upsampled result to obtain the output of the third decoder block.

[0095] The output of the third-layer decoder block is upsampled through transposed convolution, and the result is fed together with the output of the first encoder block into the AG gated attention mechanism for processing. The resulting feature map is then concatenated with the upsampled input to obtain the output of the second-layer decoder block.

[0096] The output of the second-layer decoder block is upsampled through transposed convolution, and the result, together with the output of the initial convolutional layer, is fed into the AG-gated attention mechanism for processing. The resulting feature map is then concatenated with the upsampled input to obtain the output of the first-layer decoder block. For the specific process of the AG-gated attention mechanism, please refer to process 2.4.

[0097] Finally, the output of the first-layer decoder block is used with a double convolutional layer to restore image details, resulting in a 64-array sound source distribution map after data conversion.

[0098] Step 3: Input the converted 64-array sound source distribution map into the SE-MBCNet network model to obtain the predicted sound source distribution map. The SE-MBCNet network model is trained using the converted 64-array sound source distribution map as input and the target map as the label. The structure of the SE-MBCNet network model is as follows: Figure 3 (a)-(c).

[0099] Step 3 specifically involves:

[0100] Step 3.1: Input the transformed 64-array sound source distribution map into the two-layer convolutional module of the SE-MBCNet network model for preliminary feature extraction. The two-layer convolutional module consists of two consecutive 3×3 convolution operations, each followed by a batch normalization (BN) layer and an activation function (ReLU). This aims to capture low-level features of the input image through local feature extraction and nonlinear transformation, gradually learning more complex features.

[0101] Step 3.2: The extracted preliminary features are input into the encoder module of the SE-MBCNet network model for multi-level semantic information and feature encoding of the input feature map. The encoder module consists of four MBConv layers with SE attention mechanisms and a max-pooling layer. It extracts high-level features step by step by continuously extracting image features and downsampling them. The MBConv layers with SE attention mechanisms include depthwise separable convolutions, SE attention mechanisms, and residual connections, which are used to enhance the network's feature representation ability, reduce computational cost, and improve computational efficiency. The max-pooling layers perform downsampling, reducing spatial dimensionality and increasing the receptive field, thereby reducing computational cost and extracting more abstract features.

[0102] Step 3.3: The MBConv module of the SE attention mechanism in the SE-MBCNet network model consists of depthwise separable convolutions, the SE attention mechanism, and residual connections. First, dimensionality is increased by expanding the feature dimension through 1×1 convolutions, batch normalization (BN), and the Swish activation function, increasing the number of channels. Next, depthwise separable convolutions are performed, where 3×3 depthwise convolutions, combined with batch normalization (BN) and the Swish activation function, process each channel independently to extract local features. Then, the SE channel attention mechanism is introduced, generating channel weights through pointwise convolutions to weight the feature map. Finally, 1×1 convolutions are used to reduce the weighted feature map to the number of output channels, and residual connections are added to maintain network stability and accelerate training. The expression formula for the MBConv module of the SE attention mechanism is as follows:

[0103] X exp= Swish(BN(Conv1(X)))

[0104] X dw =Swish(BN(DwConv(X) exp )))

[0105] X out =X+Conv1(SE(X) dw ))

[0106] Here, Swish is a self-gated activation function. BN represents a batch normalization layer. Conv1 represents a 1×1 convolution. DwConv represents an independent 3×3 convolution for each channel. SE represents an SE channel attention mechanism. X exp This represents the result of the dimensionality upgrade operation. X out This represents the MBConv module results of the SE attention mechanism.

[0107] The SE channel attention mechanism module comprises three stages: compression, activation, and scaling. First, a global average pooling layer is used for compression to effectively capture global features and reduce computational complexity. Then, the activation stage adaptively generates weights for each channel using a fully connected layer and a sigmoid activation function to enhance important features. Finally, the scaling stage reweights the features of each channel by element-wise multiplying the resulting channel weights with each channel of the original feature map. This process enhances the feature representation of important channels while suppressing the feature representation of unimportant channels.

[0108] Step 3.4: Input the feature maps processed by the encoder module of the SE-MBCNet network model into the decoder to recover high-resolution feature maps. The decoder is also divided into four stages. In each stage, the decoder block and the corresponding layer of the encoder directly concatenate the feature maps through skip connections, merging the downsampled and upsampled feature maps to enhance the features of the recovered image. The encoder block consists of the MBConv layer of the SE attention mechanism and an upsampling layer, which helps to preserve detailed information and avoid losing important features during the upsampling process.

[0109] Step 3.5: Finally, the final predicted sound source distribution map with sound source location coordinates and intensity information is obtained through the 1×1 convolutional layer at the end of the network.

[0110] Step 4: Perform local maxima detection on the predicted result map to obtain the sound source location and intensity.

[0111] Step 4 specifically involves:

[0112] Step 4.1: Use the local maximum detection algorithm to identify the local maxima in each row of the data matrix of the predicted sound source distribution map and find the top three maximum values.

[0113] Step 4.2: Verify whether these values ​​are true local maxima through neighborhood comparison. By comparing the target value with its neighborhood values, determine whether the value is the largest locally. If local maxima are identified, sort the detected local maxima.

[0114] Step 4.3: Based on the sorting results, use coordinate transformation techniques to map the row and column indices of the maxima to actual coordinates to calculate the location and intensity of the source.

[0115] This process integrates local maximum detection, coordinate transformation, and sorting to improve the accuracy of source localization and intensity prediction.

[0116] Figure 5 (a) is the localization result at a single sound source location [-0.31, 0.27, 0], SNR = 20 dB, f = 1000 Hz, and z = 1 m. Figure 5 (b) is the localization result at the single sound source location [0.17, 0.53, 0], SNR = 20dB, f = 2000Hz, and z = 1m. Figure 5 (c) is the localization result at a single sound source location [-0.57, -0.15, 0], SNR = 20dB, f = 6000Hz, and z = 1m.

[0117] Figure 6(a) is the localization result at the dual sound source positions [-0.21, -0.19, 0], [0.59, -0.19, 0], SNR = 20dB, f = 1000Hz, z = 1m. Figure 6 (b) shows the localization results at the dual sound source locations [0.05, -0.27, 0], [-0.25, -0.39, 0], SNR = 20 dB, f = 2000 Hz, and z = 1 m. Figure 6 (c) is the localization result at the dual sound source positions [-0.59, -0.13, 0], [0.47, -0.27, 0], SNR = 20dB, f = 6000Hz, and z = 1m.

[0118] Figure 7 (a) shows the localization results at the three sound source locations [-0.07, 0.53, 0], [0.35, 0.11, 0], [-0.05, -0.41, 0], SNR = 20 dB, f = 1000 Hz, and z = 1 m. Figure 7 (b) shows the localization results at the three sound source locations [-0.15, 0.19, 0], [-0.39, 0.05, 0], [-0.31, -0.63, 0], SNR = 20dB, f = 2000Hz, and z = 1m. Figure 7 (c) is the localization result diagram at the three sound source positions [0.53, 0.37, 0], [0.11, 0.17, 0], [0.45, -0.15, 0], SNR = 20dB, f = 6000Hz, and z = 1m.

[0119] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A sound source recognition method based on array expansion and SE-moving inverse bottleneck convolutional neural network, characterized in that, The methods include: Step 1: Obtain the sound pressure cross-spectrum matrix, and use the sound pressure cross-spectrum matrix to calculate the 18-array sound source distribution map and the 64-array sound source distribution map of the MUSIC algorithm; Step 2: Input the 18-array sound source distribution map into the EAG-U-Net data conversion model to obtain the converted 64-array sound source distribution map; the EAG-U-Net data conversion model is trained using the 18-array sound source distribution map as input and the 64-array sound source distribution map as labels. Step 3: Input the transformed 64-array sound source distribution map into the SE-MBCNet network model to obtain the predicted sound source distribution map; the SE-MBCNet network model is trained using the transformed 64-array sound source distribution map as input and the target map as the label. Obtaining the predicted sound source distribution map includes: Step 3.1: Input the 64-array sound source distribution map after data conversion into the two-layer convolution module of the SE-MBCNet network model. The two-layer convolution module is used to perform preliminary feature extraction on the 64-array sound source distribution map after data conversion. The two-layer convolution module includes two consecutive 3×3 convolution operations, batch normalization layer BN, and activation function ReLU. Step 3.2: Input the extracted preliminary features into the encoder module of the SE-MBCNet network model. The encoder module is used to encode multi-level semantic information and features into the feature map after preliminary feature extraction. The encoder module consists of four MBConv layers with SE attention mechanism and a max pooling layer. It extracts second-level features step by step by continuously extracting image features and downsampling them. The MBConv layer with SE attention mechanism includes: depthwise separable convolution, SE attention mechanism and residual connection. Step 3.3: Input the feature map processed by the encoder module of the SE-MBCNet network model into the decoder to recover the high-resolution feature map; the decoder is divided into four stages, and the decoder block of each stage directly concatenates the feature maps with the corresponding layer of the encoder through skip connections, merging the downsampled feature map and the upsampled feature map to enhance the features of the recovered image; the encoder block consists of the MBConv layer of the SE attention mechanism and the upsampling layer, which retains detailed information, thereby avoiding the loss of important features during the upsampling process; Step 3.4: Obtain the final predicted sound source distribution map with sound source location coordinates and intensity information through the 1×1 convolutional layer at the end of the network; Step 4: Perform local maxima detection on the predicted sound source distribution map to obtain the sound source location and intensity.

2. The sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network according to claim 1, characterized in that, Obtaining the sound source distribution map of the 18 array and the sound source distribution map of the 64 array includes: Step 1.1: Mesh the sound source plane, discretize the sound source calculation plane to form a mesh surface, i.e., the focusing plane, and take each mesh point of the focusing plane as a potential sound source point; Step 1.2: Arrange multiple microphone sensors in the sound field formed by the radiation of several sound sources to form a planar microphone array, i.e., a measurement plane, and make the measurement plane parallel to the sound source plane; Step 1.3: Receive frequency sound pressure signal data collected by the sensor array, wherein the sound pressure signal data includes: sound source signal and noise signal; Step 1.4: Preprocess the frequency sound pressure signal data, calculate the sound pressure cross spectrum matrix, and obtain the 18-array sound source distribution map and the 64-array sound source distribution map according to the MUSIC algorithm.

3. The sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network according to claim 1, characterized in that, Obtaining the converted 64-array sound source distribution map includes: Step 2.1: Input the 18-array sound source distribution map into the initial convolutional layer of the EAG-U-Net data conversion model. The initial convolutional layer uses two 3×3 convolutions to perform preliminary processing on the input image and extract the basic features of the 18-array sound source distribution map. Step 2.2: Input the basic features into the encoder of the EAG-U-Net data transformation model. The encoder consists of three layers: a double convolutional layer and a pooling layer. It extracts image features continuously and performs downsampling to gradually extract the first high-level features. The double convolutional layer contains two 3×3 convolutions, a ReLU activation function, and a batch normalization layer. The pooling layer uses max pooling for downsampling to reduce the spatial dimension and increase the receptive field, thereby reducing the computational load and extracting more abstract features, namely the first high-level features. Step 2.3: After the pooling operation in the third encoder block, an ECA attention mechanism is added. The ECA attention mechanism obtains the average value of each channel by performing global average pooling to capture global information. Based on the global information, a one-dimensional convolution operation is performed on the pooled channel information to capture the dependency relationship between channels. Based on the dependency relationship, a weight coefficient is obtained through the Sigmoid activation function. The weight coefficient is used to perform a weighted operation on each channel. Step 2.4: Using the AG gating attention mechanism, the encoder feature map is... and the corresponding level of decoder feature map The features are concatenated to generate a first fused feature map. A 1×1 convolution is then applied to this first fused feature map to extract higher-level feature information, reducing the number of channels and generating input features for an attention map. An attention weight map A is generated using a Sigmoid activation function. This attention weight map A is then combined with the encoder feature map. Element-wise multiplication is used to apply attention weights to the feature map of the encoder, dynamically adjusting the weights of the feature map. The weighted feature map is then fused with the feature map in the decoder to obtain the second fused feature map. When restoring spatial resolution, the decoder concatenates the feature map of the encoder with the feature map of the encoder to form a first fused feature map. The first fused feature map contains the low-level features of the encoder and the high-level semantic information of the decoder. Step 2.5: Upsample the output of the third encoder block through transposed convolution. The result and the output of the second encoder block are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the third decoder block. The output of the third-layer decoder block is upsampled through transposed convolution. The result and the output of the first encoder block are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the second-layer decoder block. The output of the second layer decoder block is upsampled through transposed convolution. The result and the output of the initial convolutional layer are both input into the AG gated attention mechanism for processing. The resulting feature map is concatenated with the input upsampled result to obtain the output of the first layer decoder block. The output of the first layer decoder block is used to restore image details using a double convolutional layer to obtain a 64-array sound source distribution map after data conversion.

4. The sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network according to claim 3, characterized in that, The workflow of the depthwise separable convolution, SE attention mechanism, and residuals includes: First, dimensionality is increased by expanding the feature dimension using 1×1 convolutions, batch normalization (BN), and the Swish activation function, thereby increasing the number of channels. Next, depthwise separable convolutions are performed, where 3×3 depthwise convolutions, combined with BN and the Swish activation function, process each channel independently to extract local features. Subsequently, a channel attention (SE) mechanism is introduced, first using a global average pooling layer for compression to effectively capture global features and reduce computational complexity. Then, in the activation phase, fully connected layers and the Sigmoid activation function adaptively generate weights for each channel to enhance important features. In the scaling phase, the features of each channel are reweighted by multiplying the obtained channel weights element-wise with each channel of the original feature map. Finally, 1×1 convolutions are used to reduce the weighted feature map to the number of output channels, and residual connections are added to maintain network stability and accelerate training.

5. The sound source recognition method based on array expansion and SE moving inverse bottleneck convolutional neural network according to claim 1, characterized in that, Obtaining the sound source location and intensity includes: Step 4.1: Using a local maximum detection algorithm, identify the local maxima in each row of the data matrix of the predicted sound source distribution map and find the top three maximum values; Step 4.2: Verify whether the maximum value is a true local maximum by comparing the target value with the values ​​in its neighborhood to determine whether the maximum value is the largest in the local range. If the maximum value is the largest in the local range, sort the detected local maxima. Step 4.3: Based on the sorting results, use coordinate transformation technology to map the row and column indices of the maximum values ​​to the actual coordinates to calculate the location and intensity of the source.

Citation Information

Patent Citations

  • Arc and corona discharge detection device based on microphone array voiceprint feature positioning algorithm

    CN118625063A

  • Method and device of channel equalization and beam controlling for a digital speaker array system

    US20130108078A1