A sound source direction recognition method based on three-element micro-microphone array
By using a three-element microphone array and a lightweight neural network RepMobileViT, the problem of sound source localization in a reverberant environment using a small-size microphone array is solved, efficient and lightweight sound source orientation estimation is achieved, and the accuracy and robustness of sound source localization are improved.
Patent Information
- Application Number
- CN202411579883.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-07
AI Technical Summary
In indoor environments, traditional sound source localization methods are severely affected by reverberation, and the sound intensity characteristics are complex under small-size microphone arrays, lacking lightweight and real-time performance. Existing deep learning methods have failed to effectively solve the problem of sound source orientation estimation in small-size microphone arrays.
A sound source orientation recognition method based on a three-element micro-microphone array is designed. By constructing RGB sound intensity spectrogram and RGB masked sound intensity spectrogram, and combining them with a lightweight neural network RepMobileViT for training and testing, the accuracy and robustness of sound source localization are improved.
It achieves efficient and lightweight sound source direction estimation in reverberant environments, with the resolution increased to 5 degrees, improving the accuracy and timeliness of sound source positioning, and is suitable for small-size microphone arrays.
Smart Images

Figure CN119170043B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sound source localization, and in particular is a sound source orientation recognition method based on a three-element miniature microphone array. Background Art
[0002] Sound source localization plays a crucial role in modern communication and sensing systems, such as navigation, human-computer interaction, rescue operations, and intelligent surveillance. In practical applications, sound source localization is often simplified to direction-of-arrival (DOA) estimation. Traditional sound source direction estimation methods include: first, time difference of arrival (TDOA)-based DOA estimation methods, such as phase-shifted generalized cross-correlation (GCC-PHAT); second, beamforming methods, such as phase-shifted steerable response power (SRP-PHAT); and third, high-resolution spectral estimation techniques, such as multiple signal classification (MUSIC) and signal parameter estimation via rotation-invariant techniques (ESPRIT). However, in indoor environments, the presence of reverberation can significantly degrade the performance of these traditional methods. To address this issue, a feasible approach is to utilize larger arrays and an expanded number of microphones, as larger arrays imply higher spatial diversity. However, in some practical applications, such as mobile devices and confined vehicle space, the number of microphones may be limited, necessitating the preference for smaller array configurations. This poses a challenge to traditional sound source direction estimation methods.
[0003] At the same time, the sound source estimation method based on sound intensity can estimate the sound intensity related to the sound source direction estimation by simultaneously measuring the sound pressure and particle velocity, which is very beneficial for accurately estimating DOA. Its core is to use the finite difference of the pressure measurement values of two adjacent microphones to approximate the pressure gradient and construct it into a first-order differential microphone array DMA. Due to the small size of the first-order DMA array itself, it provides a good solution for sound source direction estimation under small-size arrays. At present, there have been related research and patent inventions on basic differential microphone small-size arrays and sound source direction estimation based on sound intensity, but overall the sound intensity characteristics obtained are relatively complex and lack certain lightweight and real-time considerations.
[0004] In recent years, data-driven deep learning techniques have shown great potential in sound source localization. They can learn the nonlinear relationship between the information contained in acoustic features and the location of sound sources, thereby improving the accuracy and robustness of sound source localization. Therefore, combining sound source direction estimation based on differential microphone intensity characteristics with deep learning techniques is conducive to the development of sound source localization using small-sized microphone arrays in reverberant environments, and is currently a hot topic in fields such as artificial intelligence. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a sound source orientation recognition method based on a three-element miniature microphone array. The signals received by the three-element microarray are converted into a color RGB sound intensity spectrogram to produce a data set, and a dedicated lightweight neural network is designed for training and testing the effectiveness of sound source estimation.
[0006] The technical solution of the present invention is:
[0007] Step 1) designing a microphone array having three microphone units and performing sound intensity preprocessing on the received signal;
[0008] Step 2) Using a differential microphone group formed by two right-angled sides of a microphone array having an isosceles right triangle structure, obtain the sound intensities of the H and R axes. The sound intensities of the H and R axes are placed in the Green and Blue channels of the RGB color image channels, respectively. The combined sound intensity spectrogram is formed into an RGB sound intensity spectrogram. A small-value masking layer is designed to obtain a masked sound intensity spectrogram. The masked sound intensity grayscale image of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram to generate an RGB masked sound intensity spectrogram.
[0009] Step 3) The horizontal plane is divided into 72 angle categories at fixed intervals, and an RGB sound intensity spectrogram dataset and an RGB masked sound intensity spectrogram dataset are generated for each category;
[0010] Step 4) Construct a lightweight neural network RepMobileViT;
[0011] Step 5) Feed the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset into the neural network for training and testing to evaluate the sound source localization performance.
[0012] Furthermore, step 1) is specifically as follows:
[0013] 101) Construct a three-element microphone array with an isosceles right triangle structure. The three microphone units 1, 2, and 3 are denoted as M1, M2, and M3, respectively, and are located at the three vertices of the triangle. Microphone 2 is located at the right-angled vertex of the triangle. The distance between microphones 1 and 2 is denoted as the array diameter, which is 4 cm.
[0014] 102) multiplexing the right-angle vertex microphone No. 2, forming a differential microphone array with the direction of the H axis of microphones No. 1 and No. 2, and forming a differential microphone array with the direction of the R axis of microphones No. 3 and No. 2, to obtain two orthogonal microphone groups;
[0015] 103) Perform signal preprocessing to obtain the sound pressure and vibration velocity related to the sound source information:
[0016]
[0017] (1) Where V H (ω,t) represents the vibration velocity component of each time-frequency point in the H-axis direction where microphone No. 1 and microphone No. 2 are located. (2) Where V R (ω,t) represents the vibration velocity component of each time-frequency point in the R-axis direction where microphone No. 3 and microphone No. 2 are located, P i (ω, t) is the short-time Fourier transform of the sound pressure at microphone i, i = 1, 2, 3, (ω, t) represents the time-frequency point, j represents an imaginary number, ρ represents the air density, and d represents the array size of 4 cm;
[0018] 104) The sound intensity component at each time-frequency point is obtained by using the sound pressure and vibration velocity:
[0019]
[0020]
[0021] (3) Where I H (ω,t) is the component of the sound intensity at each time-frequency point in the H-axis direction at the coordinate of the upper microphone No. 2, (4) where I R (ω,t) is the component of the sound intensity at each time-frequency point in the R-axis direction at the coordinate of the upper microphone No. 2, and Re{·} represents the complex real part;
[0022] 105) Preprocess the obtained sound intensity components, normalize them, and map the sound intensity values to the range of [0, 255]:
[0023]
[0024] (5) represents the normalized component of the H-axis sound intensity component, where Represents the normalized component of the H-axis sound intensity component;
[0025]
[0026] (7) represents the sound intensity grayscale image of the H-axis sound intensity component, (8) where Represents the grayscale image component of the R-axis sound intensity component, which is used to construct the layer of the RGB sound intensity spectrogram.
[0027] Furthermore, step 2) is specifically as follows:
[0028] 201) Normalize and map the sound intensity component to the value range of [0,255] and The Green and Blue channels placed on the RGB color image channels, referred to as the G and B channels, are combined to form an RGB sound intensity spectrogram;
[0029] 202) Furthermore, in order to improve the stability of sound source localization of RGB sound intensity spectrogram in reverberant noise environment, a small value masking layer is designed and a binary masking function is introduced to refine the selection of reliable TF points:
[0030]
[0031] (9) where B(ω,t) is the binary masking function, I * (ω,t) represents the sound intensity component of the * axis, I * (ω,t) represents the average value of the sum of the sound intensities of all time-frequency points of the m×n size numerical matrix after the short-time Fourier transform of the sound intensity component;
[0032] 203) Binary masking is introduced into the sound intensity grayscale image, and the masked sound intensity spectrogram is expressed as:
[0033]
[0034] (10) Represents the grayscale image of the masking sound intensity on the H axis;
[0035] 204) Based on the RGB sound intensity spectrogram, the masking sound intensity grayscale image of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram, referred to as the R channel, to generate the RGB masking sound intensity spectrogram.
[0036] Furthermore, step 3) is specifically as follows:
[0037] 301) The 360-degree horizontal plane is divided into 72 angle categories with an interval of 5 degrees as the positioning angle categories for sound source direction estimation. A speech dataset totaling 300 1-second speech sounds is selected. The RIR generator software is used to generate array signals after the room impulse response at each angle from the original clean speech. The parameters of the simulated RIR impulse response are configured to control the signal-to-noise ratio (SNR) in the range of 5dB to 30dB, the reverberation time (RT60) in the range of 0.2s to 1.0s, and the room size in the range of 7×6×3 (m). To enrich the dataset and improve the versatility of direction estimation, for each RIR room impulse response, the SNR and RT60 are randomly fluctuated within the basic range, the room size is randomly varied by ±1m from the basic size, and the sound source position is randomly set at 1m, 2m, and 3m.
[0038] 302) generating RGB sound intensity spectrograms and RGB masking sound intensity spectrograms corresponding to 72 sound source azimuth angles for the simulated array signal;
[0039] 303) Completed the dataset production, a total of two spectrogram datasets, one is RGB sound intensity spectrogram, the other is RGB masked sound intensity spectrogram, for each spectrogram is produced into a sound source localization dataset with 72 angle categories; finally, 300 voices × 3 sound source distances × 72 angle categories × 2 spectrograms were obtained, totaling 129,600 images as the dataset.
[0040] Furthermore, step 4) is specifically as follows:
[0041] 401) Design a network feature extraction layer based on RepViT to enhance the local feature extraction capability of neural networks at different resolutions. At the same time, introduce the MetaFormer structure of RepViT to maintain efficient local feature extraction while capturing global information.
[0042] 402) The downsampling layer with the RepViT structure is fused with the input of the MobileViT network module to form a new lightweight neural network learning module, which is recorded as a RepMobileViT module;
[0043] 403) Reference Figure 6 By connecting RepViT and RepMobileViT modules in series and configuring the number and structure of modules, the designed lightweight network RepMobileViT network is formed; specifically, referring to Figure 7 The detailed submodules of the complete network are shown, namely: (a) RepViT block, which contains a DW convolution, SE module and 1×1 convolution; (b) RepViT downsampling block, which contains a RepViT block core, downsampling DW convolution with a stride of 2, 1×1 convolution and FFN module, where B represents the batch size of the input feature, C1 represents the number of input channels, H1 represents the height of the input feature, and W1 represents the width of the input feature; (c) MobileViT-v3 module, which contains a local representation block composed of DW convolution, a global representation block composed of Transformer and a feature fusion module, where L represents the number of Transformer blocks, h and w represent the size of the feature block are both 2, Cin represents the number of channels of the input feature, Cout represents the number of channels of the output feature, H represents the height of the input feature, and W represents the width of the input feature; use Figure 7 The three submodules in the neural network are designed to achieve Figure 6The complete neural network shown in the figure has five core modules implemented: Layer 1 contains a RepViT module; Layer 2 contains a RepViT downsampling block and two RepViT blocks; Layer 3 contains a RepViT downsampling block and a MobileViT-v3 module; Layer 4 contains a RepViT downsampling block and a MobileViT-v3 block; and Layer 5 contains a RepViT downsampling block and a MobileViT-v3 block. Finally, the network outputs the logits through a 1×1 convolution, a global pooling layer, and a fully connected linear layer.
[0044] Furthermore, step 5) is specifically as follows:
[0045] 501) performing neural network training on the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset respectively, and dividing the datasets into a training set and a test set according to an 8:2 ratio;
[0046] 502) PyTorch is used as the deep learning framework to build and train the RepMobileViT model. The specific equipment and configuration are as follows: Python 3.8 and PyTorch 1.10.1+cuda11.3; the experimental infrastructure is CPU: Intel i5-13490F; GPU: NVIDIA GeForce RTX 3060Ti; In addition, the Label Smoothing Loss Cross-Entropy Loss function and AdamW optimizer are used to update the model parameters; the cosine annealing learning rate scheduler is used, and the initial learning rate is 0.0125;
[0047] 503) The training set data set and its category labels are fed into the neural network for training for 200 rounds, and the weight file best.pt of the best training result is taken as the final weight for testing;
[0048] 504) Use best.pt to test and verify the test set and compare the corresponding indicators; calculate the mean absolute error, accuracy and φ-accuracy to evaluate the performance of the proposed method:
[0049]
[0050] (11) where θ i represents the actual angle of the sound source, Represents the estimated angle output by the neural network, and N represents the number of angle estimates;
[0051]
[0052] (12) where N T Indicates the total number of sound source direction evaluations, N fine It is the total number of correct sound source direction assessments; φ degree-accuracy indicates the accuracy rate when the DOA test error is within φ degrees;
[0053] 505) Using the RepMobileViT model, the performance of RGB sound intensity spectrogram and RGB masked sound intensity spectrogram were tested respectively, and the performance under different reverberation and different signal-to-noise ratios in the simulated environment was evaluated respectively;
[0054] 506) Further using the display microphone array and its supporting equipment, the horizontal plane is divided into 12 angle categories with an interval of 30 degrees with the microphone array as the center, the real environment voice signal is recorded, and the sound source direction estimation performance in the real environment is evaluated.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] (1) The present invention is a method for identifying the direction of a sound source based on a three-element micro-microphone array. This method extracts sound intensity information related to sound source estimation and designs an RGB sound intensity spectrogram and an RGB masked sound intensity spectrogram. Compared with previous research and patents, the number of microphones in a four-element microphone array in this field is reduced to three, while also reducing the complexity of the sound intensity characteristics and the number of channels, greatly improving the practicability of the method model.
[0057] (2) This paper constructs a new lightweight neural network RepMobileViT, which provides an efficient and lightweight solution, improving the timeliness and lightweightness of the deep learning sound source direction estimation method from the perspective of network design and optimization;
[0058] (3) Compared with the previous deep learning classification methods for sound source direction estimation and sound intensity source localization, the present invention further improves the resolution to 5 degrees, that is, 72 angle categories are realized in a 360-degree plane, which can realize sound source direction estimation in a high-reverberation room. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Schematic diagram of the system flow of the present invention;
[0060] Figure 2 A conventional four-element microphone array (left) and the three-element microphone array of the present invention (right);
[0061] Figure 3 The RGB sound intensity spectrogram design method proposed by the present invention;
[0062] Figure 4The RGB masking sound intensity spectrogram design method proposed by the present invention;
[0063] Figure 5 The RGB spectrograms of the present invention at different angles are shown as follows: (a) 0 degrees; (b) 45 degrees; (c) 90 degrees; and (d) 135 degrees. The left side of each group shows the RGB sound intensity spectrogram, and the right side shows the RGB masked sound intensity spectrogram.
[0064] Figure 6 This is a schematic diagram of the structure of the lightweight neural network RepMobileViT proposed in this invention;
[0065] Figure 7 Modules used within the lightweight network: (a) RepViT module; (b) RepViT downsampling module; (c) MobileViT-v3 module;
[0066] Figure 8 The sound source direction estimation accuracy of the proposed method and the commonly used methods in related fields for different reverberation times;
[0067] Figure 9 Figure 2. Sound source direction estimation accuracy of the proposed method and commonly used methods in related fields for different signal-to-noise ratios.
[0068] Figure 10 Schematic diagram of the test room. DETAILED DESCRIPTION
[0069] This paper discloses a sound source location identification method based on a three-element micro-microphone array, designed to address the problem of sound source localization in reverberant and spatially confined indoor environments. Furthermore, this method improves the accuracy and timeliness of sound source location estimation in reverberant environments by introducing an RGB sound intensity spectrogram and a lightweight neural network.
[0070] To help those skilled in the art better understand the present invention, this section further describes the present invention in detail with reference to the accompanying drawings and specific implementation methods. It should be noted that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0071] Specifically, refer to Figure 1 As shown, the sound source direction recognition method based on a three-element micro-microphone array includes the following steps:
[0072] Step 1) Designing a microphone array having three microphone units and performing sound intensity preprocessing on the received signal, including the following steps:
[0073] Step 101) Refer to Figure 2 (Left) Based on the four-element differential microphone array commonly used in the field, a Figure 2 (Right) A three-element microphone array has an isosceles right triangle structure. Microphones M1, M2, and M3 represent microphones 1, 2, and 3, respectively, and are located at the three vertices of the triangle. Microphone M2, number 2, is located at the right-angled vertex of the triangle. The distance between microphones 1 and 2 is denoted as the array diameter, which is 4 cm.
[0074] Step 102) Based on the XY coordinate system, an H-axis pointing from microphone 2 to microphone 1 and an R-axis pointing from microphone 2 to microphone 3 are constructed; using the right-angle vertex microphone 2, a differential microphone array is formed with the H-axis in the direction of microphones 1 and 2, and a differential microphone array is formed with the R-axis in the direction of microphones 3 and 2, thereby obtaining two orthogonal microphone groups;
[0075] Step 103) Perform signal preprocessing and convert it into a two-dimensional expression in the time-frequency domain through short-time Fourier transform (STFT) to obtain the sound pressure and vibration velocity related to the sound source information:
[0076]
[0077] (1) Where V H (ω,t) represents the vibration velocity component of each time-frequency point in the H-axis direction where microphone No. 1 and microphone No. 2 are located. (2) Where V R (ω,t) represents the vibration velocity component of each time-frequency point in the R-axis direction where microphone No. 3 and microphone No. 2 are located, P i (ω, t) is the short-time Fourier transform of the sound pressure at microphone i, i = 1, 2, 3, (ω, t) represents the time-frequency point, j represents an imaginary number, ρ represents the air density, and d represents the array size of 4 cm;
[0078] Step 104) Obtain the sound intensity component at each time-frequency point using the sound pressure and vibration velocity:
[0079]
[0080] (3) Where I H (ω,t) is the component of the sound intensity at each time-frequency point in the H-axis direction at the coordinate of the upper microphone No. 2, (4) where I R (ω,t) is the component of the sound intensity at each time-frequency point in the R-axis direction at the coordinate of the upper microphone 2, and Re{·} represents the complex real part. Based on the sound intensity components of the H-axis and the R-axis, a pair of orthogonal sound intensity features can be obtained for subsequent preprocessing.
[0081] Step 105) pre-process the obtained sound intensity components, normalize them and map the sound intensity values to the range of [0, 255]:
[0082]
[0083] (5) represents the normalized component of the H-axis sound intensity component, where Represents the normalized component of the H-axis sound intensity component;
[0084]
[0085] (7) represents the sound intensity grayscale image of the H-axis sound intensity component, (8) where Represents the grayscale image component of the R-axis sound intensity component, which is used to construct the layer of the RGB sound intensity spectrogram.
[0086] Step 2) extracting the RGB sound intensity spectrogram, comprising the following steps:
[0087] Step 201) Refer to Figure 3 , which is normalized and mapped to the sound intensity component in the range of [0,255] and Placed in the Green channel and Blue channel on the RGB color image channel, referred to as the G channel and R channel, combined into an RGB sound intensity spectrogram;
[0088] Step 202) Further, to improve the sound source localization stability of the RGB sound intensity spectrogram in a reverberant noise environment, a small value masking layer is designed, and a binary masking function is introduced to refine the selection of reliable TF points:
[0089]
[0090] (9) where B(ω,t) is the binary masking function, I * (ω,t) represents the sound intensity component of the * axis, I * (ω,t) represents the average value of the sum of the sound intensities of all time-frequency points of the m×n size numerical matrix after the short-time Fourier transform of the sound intensity component;
[0091] Step 203) introduces binary masking into the sound intensity grayscale image, and the masked sound intensity spectrogram is represented as:
[0092]
[0093] (10) Represents the grayscale image of the masking sound intensity on the H axis;
[0094] Step 204) Refer to Figure 4 On the basis of the RGB sound intensity spectrogram, the masking sound intensity grayscale map of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram, referred to as the R channel, to generate the RGB masking sound intensity spectrogram.
[0095] Step 3) Create a DOA label dataset, including the following steps:
[0096] Step 301) The 360-degree horizontal plane is divided into 72 angle categories with an interval of 5 degrees as the positioning angle categories for sound source direction estimation. A speech data set totaling 300 1-second speech sounds is selected, and the RIR generator software is used to generate an array signal after the room impulse response at each angle is generated from the original clean speech. The parameters of the simulated RIR impulse response are configured to control the signal-to-noise ratio (SNR) within the range of 5 dB to 30 dB, the reverberation time (RT60) within the range of 0.2 s to 1.0 s, and the room size within the range of 7 × 6 × 3 m. To enrich the data set and improve the versatility of direction estimation, for each RIR room impulse response, the SNR and RT60 are randomly fluctuated within a basic range, the room size is randomly varied by ±1 m from the basic size, and the sound source position is randomly set between 1 m, 2 m, and 3 m.
[0097] Step 302) Generate RGB sound intensity spectrograms and RGB masking sound intensity spectrograms corresponding to 72 sound source azimuth angles for the simulated array signal, refer to Figure 5 , shows the RGB sound intensity spectrogram (left) and RGB masked sound intensity spectrogram (right) at 0 degrees (a), 45 degrees (b), 90 degrees (c), and 135 degrees (d);
[0098] Step 303) Complete the dataset production. There are two types of spectrogram datasets in total: one is an RGB sound intensity spectrogram and the other is an RGB masked sound intensity spectrogram. For each spectrogram, a sound source localization dataset with 72 angular categories is produced. Finally, 300 voices × 3 sound source distances × 72 angular categories × 2 spectrograms are obtained, totaling 129,600 images as the dataset.
[0099] Step 4) Design a lightweight neural network, including the following steps:
[0100] Step 401) Designing a network feature extraction layer based on RepViT to enhance the local feature extraction capability of the neural network at different resolutions, while introducing the MetaFormer structure of RepViT to maintain efficient local feature extraction while capturing global information;
[0101] Step 402) The downsampling layer with the RepViT structure is fused with the input of the MobileViT network module to form a new lightweight neural network learning module, denoted as the RepMobileViT module;
[0102] Step 403) Refer to Figure 6 By connecting RepViT and RepMobileViT modules in series and configuring the number and structure of modules, the designed lightweight network RepMobileViT network is formed. Figure 7 The detailed submodules of the complete network are shown, namely: (a) RepViT block, which contains a DW convolution, SE module and 1×1 convolution; (b) RepViT downsampling block, which contains a RepViT block core, downsampling DW convolution with a stride of 2, 1×1 convolution and FFN module, where B represents the batch size of the input feature, C1 represents the number of input channels, H1 represents the height of the input feature, and W1 represents the width of the input feature; (c) MobileViT-v3 module, which contains a local representation block composed of DW convolution, a global representation block composed of Transformer and a feature fusion module, where L represents the number of Transformer blocks, h and w represent the size of the feature block are both 2, Cin represents the number of channels of the input feature, Cout represents the number of channels of the output feature, H represents the height of the input feature, and W represents the width of the input feature. Use Figure 7 The three submodules in the neural network are designed to achieve Figure 6 The complete neural network shown in the figure has five core modules implemented: Layer 1 contains a RepViT module; Layer 2 contains a RepViT downsampling block and two RepViT blocks; Layer 3 contains a RepViT downsampling block and a MobileViT-v3 module; Layer 4 contains a RepViT downsampling block and a MobileViT-v3 block; and Layer 5 contains a RepViT downsampling block and a MobileViT-v3 block. Finally, the network outputs the logits through a 1×1 convolution, a GlobalPool global pooling layer, and a Linear fully connected layer.
[0103] Step 5) Deep learning training and testing, including the following steps:
[0104] Step 501) performing neural network training on the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset respectively, and dividing the datasets into a training set and a test set according to an 8:2 ratio;
[0105] Step 502) PyTorch is used as a deep learning framework to build and train the RepMobileViT model. The specific equipment and configuration are as follows: Python 3.8 and PyTorch 1.10.1+cuda11.3; the experimental infrastructure is CPU: Intel i5-13490F; GPU: NVIDIA GeForce RTX 3060Ti; In addition, the Label Smoothing Loss Cross-Entropy Loss function and AdamW optimizer are used to update the model parameters; the cosine annealing learning rate scheduler is used, and the initial learning rate is 0.0125;
[0106] Step 503) The training set data set and its category labels are fed into the neural network for training for 200 rounds, and the weight file best.pt of the best training result is taken as the final weight for testing;
[0107] Step 504) Use best.pt to test and verify the test set and compare the corresponding indicators; calculate the mean absolute error, accuracy and φ-accuracy to evaluate the performance of the proposed method:
[0108]
[0109] (11) where θ i represents the actual angle of the sound source, Represents the estimated angle output by the neural network, and N represents the number of angle estimates;
[0110]
[0111] (12) where N T Indicates the total number of sound source direction evaluations, N fine It is the total number of correct sound source direction assessments; φ degree-accuracy indicates the accuracy rate when the DOA test error is within φ degrees;
[0112] Step 505) The RepMobileViT model is used to test the performance of the RGB sound intensity spectrogram and the RGB masked sound intensity spectrogram respectively, and the performance under different reverberations and different signal-to-noise ratios in the simulated environment is evaluated respectively. As a reference, the common sound source orientation estimation models in the current research field are compared, including: the GCC-PHAT feature model (GCC-PHAT-CNN) trained using CNN, the normalized phase weighted redundancy optimization sound intensity method model (SI-PNR-CNN) trained using LSSVM, the normalized phase weighted sound intensity method model (SI-CNN) trained using CNN, the whitening weighted method (SI-WW) trained using CNN, and the RGB sound intensity spectrogram method and RGB masked sound intensity spectrogram method trained using the lightweight neural network RepMobileViT network proposed in this paper; in the comparative experiment, when the signal-to-noise ratio (SNR) is fixed at 10 dB and the reverberation RT60 increases from 0.2s to 1.0s, the 5-degree Accuracy of various methods is compared with the results of the 5-degree-accuracy. Figure 8 , When the reverberation RT60 is fixed at 0.6s and the signal-to-noise ratio (SNR) increases from 5dB to 30dB, the 5-degree accuracy of various methods is compared. Figure 9 ;
[0113] Step 506) The microphone array and its supporting equipment are further used to divide the horizontal plane into 12 angle categories with an interval of 30 degrees with the microphone array as the center. The room layout is referenced Figure 10 , record 20 segments of 0.5 seconds of audio for each angle of the real environment speech signal, evaluate the sound source direction estimation performance in the real environment, and refer to Table 1 for the accuracy of the model.
[0114] Table 1
[0115]
[0116] The present invention extracts sound intensity information from a three-element microphone array, proposes RGB sound intensity spectrograms and RGB masked sound intensity spectrograms, and designs a lightweight neural network to solve the problem of sound source direction estimation. The method includes signal sound intensity preprocessing, RGB sound intensity spectrogram design, preparation of a sound source direction estimation dataset, lightweight neural network design, and deep learning training and testing. Ultimately, the complete system can extract the RGB sound intensity spectrogram from the array received signal and derive the direction estimate of the sound source through the designed neural network. It achieves good sound source localization accuracy in different reverberant and noisy environments. The accuracy in real environments further demonstrates that the RGB masked sound intensity spectrogram has higher accuracy and robustness than the RGB sound intensity spectrogram in real environments. Compared with previous classic algorithms in related fields, the method proposed by the present invention has a lighter feature map design and a lighter and more efficient network design. It can greatly reduce feature calculations while maintaining a high positioning accuracy, reflecting the innovation and efficiency of the present invention.
[0117] It should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A sound source direction recognition method based on a three-element micro-microphone array, which extracts the sound intensity information of the signal and converts it into a sound intensity spectrogram, and then feeds it into a neural network for training and testing to achieve sound source direction estimation, characterized by The following steps are involved: Step 1) designing a microphone array having three microphone units and performing sound intensity preprocessing on the received signal; Step 2) Using a differential microphone group formed by two right-angled sides of a microphone array having an isosceles right triangle structure, obtain the sound intensities of the H and R axes. The sound intensities of the H and R axes are placed in the Green and Blue channels of the RGB color image channels, respectively. The combined sound intensity spectrogram is formed into an RGB sound intensity spectrogram. A small-value masking layer is designed to obtain a masked sound intensity spectrogram. The masked sound intensity grayscale image of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram to generate an RGB masked sound intensity spectrogram. Step 3) The horizontal plane is divided into 72 angle categories at fixed intervals, and an RGB sound intensity spectrogram dataset and an RGB masked sound intensity spectrogram dataset are generated for each category; Step 4) Construct a lightweight neural network RepMobileViT; Step 5) Feed the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset into the neural network for training and testing to evaluate the sound source localization performance.
2. A method for identifying the direction of a sound source based on a three-element micro-microphone array according to claim 1, characterized in that The steps of designing a microphone array having three microphone units and performing sound intensity preprocessing on the received signal in step 1 are as follows: 201) Construct a three-element microphone array with an isosceles right triangle structure. The three microphone units 1, 2, and 3 are denoted as M1, M2, and M3, respectively, and are located at the three vertices of the triangle. Microphone 2 is located at the right-angled vertex of the triangle. The hypotenuse of the triangle is denoted as the microphone array. The diameter of the microphone array is 4 cm. 202) multiplexing the right-angle vertex microphone No. 2, forming a differential microphone array with the direction of the H axis of microphones No. 1 and No. 2, and forming a differential microphone array with the direction of the R axis of microphones No. 3 and No. 2, to obtain two orthogonal microphone groups; 203) Perform signal preprocessing to obtain the sound pressure and vibration velocity related to the sound source information: In the above formula, V H (ω,t) represents the vibration velocity component of each time-frequency point in the H-axis direction where microphone 1 and microphone 2 are located, V R (ω,t) represents the vibration velocity component of each time-frequency point in the R-axis direction where microphone No. 3 and microphone No. 2 are located, P i (ω, t) is the short-time Fourier transform of the sound pressure at microphone i, i = 1, 2, 3, (ω, t) represents the time-frequency point, j represents an imaginary number, ρ represents the air density, and d represents the array size of 4 cm; 204) The sound intensity component at each time-frequency point is obtained by using the sound pressure and vibration velocity: In the above formula, I H (ω,t) is the component of the sound intensity at each time-frequency point in the H-axis direction at the coordinate of the upper microphone No. 2, I R (ω,t) is the component of the sound intensity at each time-frequency point in the R-axis direction at the coordinate of the upper microphone No. 2, and Re{·} represents the complex real part; 205) The obtained sound intensity components are preprocessed, normalized, and the sound intensity values are mapped to the range of [0, 255]: In the above formula, represents the normalized component of the H-axis sound intensity component, Represents the normalized component of the H-axis sound intensity component; In the above formula, The sound intensity grayscale image representing the H-axis sound intensity component, Represents the grayscale image component of the R-axis sound intensity component, which is used to construct the layer of the RGB sound intensity spectrogram.
3. The method for identifying the direction of a sound source based on a three-element micro-microphone array according to claim 1, characterized in that In step 2, the sound intensities of the H and R axes are acquired by using a differential microphone group formed by two right-angled sides in a microphone array having an array structure of an isosceles right triangle. The sound intensities of the H and R axes are respectively placed in the Green and Blue channels of the RGB color image channels. After combination, an RGB sound intensity spectrogram is formed. A small-value masking layer is designed to obtain a masked sound intensity spectrogram. The masked sound intensity grayscale image of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram to generate an RGB masked sound intensity spectrogram. The specific steps are as follows: 301) Normalize and map the sound intensity component to the value range of [0,255] and The Green and Blue channels placed on the RGB color image channels, referred to as the G and B channels, are combined to form an RGB sound intensity spectrogram; 302) Furthermore, in order to improve the stability of sound source localization of RGB sound intensity spectrogram in reverberant noise environment, a small value masking layer is designed and a binary masking function is introduced to refine the selection of reliable TF points: In the above formula, B(ω,t) is a binary masking function, I * (ω,t) represents the sound intensity component of the * axis, Represents the average value of the sum of the sound intensities of all time-frequency points of the m×n size numerical matrix after the short-time Fourier transform of the sound intensity component; 303) Binary masking is introduced into the sound intensity grayscale image, and the masked sound intensity spectrogram is expressed as: In the above formula, Represents the grayscale image of the masking sound intensity on the H axis; 304) Based on the RGB sound intensity spectrogram, the masking sound intensity grayscale image of the H axis is selected and placed in the Red channel of the RGB sound intensity spectrogram, referred to as the R channel, to generate the RGB masking sound intensity spectrogram.
4. The method for identifying the direction of a sound source based on a three-element micro-microphone array according to claim 1, characterized in that In step 3, the horizontal plane is divided into 72 angle categories at fixed intervals, and an RGB sound intensity spectrogram dataset and an RGB masked sound intensity spectrogram dataset are produced for each category. The specific steps are as follows: 401) Divide the 360-degree horizontal plane into 72 angle categories with an interval of 5 degrees, select a speech data set and use RIRgenerator software to generate an array signal after the room impulse response of each angle for the original clean speech; 402) generating RGB sound intensity spectrograms and RGB masking sound intensity spectrograms corresponding to 72 sound source azimuth angles for the simulated array signal; 403) Complete the dataset production, a total of two spectrogram datasets, one is RGB sound intensity spectrogram, the other is RGB masked sound intensity spectrogram, for each spectrogram is produced into a sound source localization dataset with 72 angle categories.
5. The method for identifying the direction of a sound source based on a three-element micro-microphone array according to claim 1, characterized in that The construction of the lightweight neural network RepMibileViT in step 4 is designed as follows: 501) Design a network feature extraction layer based on RepViT to enhance the local feature extraction capability of the neural network at different resolutions. At the same time, introduce the MetaFormer structure of RepViT to maintain efficient local feature extraction while capturing global information; 502) The downsampling layer with the RepViT structure is fused with the input of the MobileViT network module to form a new lightweight neural network learning module, which is recorded as the RepMobileViT module; 503) By connecting RepViT and RepMobileViT modules in series and configuring the number and configuration of modules, a designed lightweight network RepMobileViT network is formed.
6. The method for identifying the direction of a sound source based on a three-element micro-microphone array according to claim 1, characterized in that In step 5, the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset are fed into the neural network for training and testing to evaluate the sound source localization performance. The specific steps are as follows: 601) performing neural network training on the RGB sound intensity spectrogram dataset and the RGB masked sound intensity spectrogram dataset respectively, and dividing the datasets into a training set and a test set according to an 8:2 ratio; 602) Use PyTorch as a deep learning framework to build and train the RepMobileViT model; 603) The training set data set and its category labels are fed into the neural network for training for 200 rounds, and the weight file best.pt of the best training result is taken as the final weight for testing; 604) Use best.pt to test and verify the test set and compare the corresponding indicators; calculate the mean absolute error, accuracy and φ-accuracy to evaluate the performance of the proposed method: In the above formula, θ i represents the actual angle of the sound source, Represents the estimated angle output by the neural network, and N represents the number of angle estimates; In the above formula, N T Indicates the total number of sound source direction evaluations, N fine is the total number of correct sound source direction assessments; φ degree-accuracy indicates the accuracy when the DOA test error is within φ degrees; 605) Using the RepMobileViT model, the performance of RGB sound intensity spectrogram and RGB masked sound intensity spectrogram were tested respectively, and the performance under different reverberation and different signal-to-noise ratios in the simulated environment was evaluated respectively; 606) Further use the display microphone array and its supporting equipment to record real environment voice signals and evaluate the sound source direction estimation performance in the real environment.