A method, device and storage medium for sound source localization of a microphone array
By using the convolutional residual network model and subband delay difference and SRP-PHAT spatial spectrum as input, the problem of insufficient positioning performance in complex acoustic environments in the prior art is solved, and efficient and robust sound source positioning effect is achieved.
Patent Information
- Application Number
- CN202210427289.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-22
AI Technical Summary
Existing microphone array sound source positioning technology is less robust in reverberation and noise environments, making it difficult to achieve accurate positioning in complex acoustic environments.
The convolutional residual network (CRN) model is used to extract the subband time delay difference and subband SRP-PHAT spatial spectrum as spatial location cues and use it as input to the CRN model to construct the mapping relationship between the sound source orientation and the spatial location cues.
It significantly improves the performance of sound source positioning, enhances the robustness and generalization ability to noise and reverberation environments, and realizes real-time sound source positioning in complex acoustic environments.
Smart Images

Figure CN114895245B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device and storage medium for sound source localization of a microphone array, belonging to the technical field of sound source localization. Background Art
[0002] The sound source localization technology based on microphone arrays has broad application prospects and potential economic value in the front-end processing of speech recognition and speaker recognition systems, as well as in video conferencing, intelligent robots, smart homes, etc. The localization algorithms based on time delay difference and the localization algorithm based on SRP-PHAT (Steered Response Power-Phase Transform) are two typical traditional localization methods. Although these two algorithms are easy to implement, their robustness to reverberation and noise is relatively low. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies in the prior art, and provide a method, device and storage medium for sound source localization of a microphone array, which significantly improves the localization performance and has good generalization ability for unknown noise and reverberation environments.
[0004] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0005] In a first aspect, the present invention provides a method for sound source localization of a microphone array, including:
[0006] Obtain a test signal;
[0007] Preprocess the test signal to obtain a single-frame test signal;
[0008] Extract the spatial localization clues of the single-frame test signal and use it as a test sample;
[0009] Input the test sample into a pre-constructed and trained CRN model for testing, and obtain the probability that the test signal belongs to each azimuth angle. Among them, the azimuth angle with the maximum probability is taken as the azimuth angle estimation value of this frame of signal.
[0010] Further, the construction and training method of the CRN model includes:
[0011] Convolve a clean speech signal with room impulse responses at different azimuth angles, and add different degrees of noise and reverberation to generate multiple microphone array signals;
[0012] Preprocess the multiple microphone array signals to obtain multiple single-frame signals;
[0013] Extract the spatial localization cues of multiple single-frame signals, use them as the training samples of the CRN model, and at the same time mark the corresponding azimuth of each sample as the class label of the sample;
[0014] Construct a CRN model, and use the training samples and class labels as the training data set of the CRN model for training.
[0015] Further, convolve the clean speech signal with the room impulse responses at different azimuth angles, and add different levels of noise and reverberation to generate multiple microphone array signals. The formula is as follows:
[0016] x m (t) = h m (t) * s(t) + v m (t), m = 1, 2,..., M
[0017] Where, x m (t) represents the speech signal received by the m-th microphone at the specified azimuth. m is the serial number of the microphone element, m = 1, 2,..., M, and M is the number of microphone elements. s(t) is the clean speech, and h m (t) represents the room impulse response from the specified sound source azimuth to the m-th microphone. h m (t) is related to the sound source azimuth and room reverberation, and v m (t) represents noise.
[0018] Further, preprocess the multiple microphone array signals to obtain multiple single-frame signals, including:
[0019] The preprocessing includes frame division and windowing, where:
[0020] The frame division method is: use a preset frame length and frame shift to divide the time-domain signal x m (t) of the m-th element into multiple single-frame signals x m (iN + n), where i is the frame number, n represents the sampling number within a frame, 0 ≤ n < N, and N is the frame length;
[0021] The windowing method is: x m (i, n) = w H (n)x m (iN + n)
[0022] Where, x m (i, n) is the signal of the i-th frame of the m-th element after windowing processing,
[0023] is the Hamming window.
[0024] Further, the extraction of the spatial localization cues of the multiple single-frame signals includes:
[0025] Perform a discrete Fourier transform on each single-frame signal to convert the time-domain signal to a frequency-domain signal;
[0026] The calculation formula for the discrete Fourier transform is:
[0027]
[0028] where X m (i, k) is the discrete Fourier transform of x m (i, n), representing the frequency-domain signal of the m-th array element in the i-th frame, k is the frequency point, K is the length of the discrete Fourier transform, K = 2N, and DFT(·) represents the discrete Fourier transform;
[0029] Design a Gammatone filter bank, where g j (t) is the impulse response function of the j-th Gammatone filter, and its expression is:
[0030]
[0031] where j represents the filter number; C is the filter gain; t represents continuous time; a is the filter order; represents the phase; f j represents the center frequency of the j-th filter; b j represents the filter attenuation factor, and b j The calculation formula is:
[0032] b j = 1.109ERB(f j )
[0033] ERB(f j ) = 24.7(4.37f j / 1000 + 1)
[0034] Perform a discrete Fourier transform on each Gammatone filter to obtain its frequency-domain expression:
[0035]
[0036] Calculate the sub-band generalized cross-correlation function of each frame of signal, and its calculation formula is as follows:
[0037]
[0038] where R mn (i, j, τ) represents the generalized cross-correlation function of the m-th array element and the n-th array element in the i-th frame and the j-th sub-band;
[0039] Obtain the sub-band time delay difference of each frame signal, and its expression is as follows:
[0040]
[0041] Among them, T mn (i,j) represents the time delay difference between the m-th array element and the n-th array element in the j-th sub-band of the i-th frame;
[0042] Calculate the sub-band SRP-PHAT function of each frame signal, and its calculation formula is as follows:
[0043]
[0044] Among them, P(i,j,r) represents the SRP-PHAT power value of the j-th sub-band of the i-th frame signal when the beam direction of the array is r; τ mn (r) represents the time difference of sound waves propagating from the beam direction r to the m-th microphone and the n-th microphone, and its calculation formula is:
[0045]
[0046] Among them, r represents the coordinate of the beam direction, r m represents the position coordinate of the m-th microphone, c is the speed of sound in the air, f s is the signal sampling rate;
[0047] Assume that the sound source and the microphone array are on the same horizontal plane, and the sound source is in the far field of the array, then the equivalent calculation formula of τ mn (r) is:
[0048]
[0049] Among them, ξ = [cosθ, sinθ] T , θ is the azimuth angle of the beam direction r, τ mn (r) has nothing to do with the received signal, so it can be calculated offline and saved in memory;
[0050] Perform normalization processing on the sub-band SRP-PHAT function, and its calculation formula is as follows:
[0051]
[0052] Combine the time delay differences and SRP-PHAT functions of all sub-bands within the same frame to form a feature matrix, and obtain the spatial clue of the mixed feature, and the expression is as follows:
[0053]
[0054] Among them, y train (i) represents the spatial localization clue of the i-th frame signal, and J is the number of sub-bands.
[0055] Further, the method of preprocessing the test signal to obtain a single-frame test signal is the same as the method of preprocessing the multiple microphone array signals to obtain multiple single-frame signals;
[0056] The method of extracting the spatial localization clues of the single-frame test signal is the same as the method of extracting the spatial localization clues of the multiple single-frame signals.
[0057] Further, the CRN model includes an input layer, two residual blocks, a pooling layer, two fully connected layers, and a final output layer. Among them, the residual block is composed of multiple convolutional layers and batch normalization layers. Each residual block structure contains two batch normalization layers and two convolutional layers, and processes the input in the order of first batch normalization layer, then ReLU, and finally convolutional layer. Among them, the output layer uses Softmax, and the loss function is the cross-entropy function.
[0058] In a second aspect, the present invention provides a microphone array sound source localization device, including:
[0059] An acquisition unit for acquiring a test signal;
[0060] A preprocessing unit for preprocessing the test signal to obtain a single-frame test signal;
[0061] An extraction unit for extracting the spatial localization clues of the single-frame test signal and using it as a test sample;
[0062] A test unit for inputting the test sample into a pre-constructed and trained CRN model for testing, and obtaining the probability that the test signal belongs to each azimuth angle. Among them, the azimuth angle with the maximum probability is taken as the azimuth angle estimation value of this frame of signal.
[0063] In a third aspect, the present invention provides a microphone array sound source localization device, including a processor and a storage medium;
[0064] The storage medium is used to store instructions;
[0065] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the foregoing.
[0066] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of the foregoing.
[0067] Compared with the prior art, the beneficial effects achieved by the present invention:
[0068] The present invention uses sub-band time delay difference and sub-band SRP-PHAT spatial spectrum as spatial positioning cues, which have strong robustness and spatial information representation ability; the present invention uses a convolutional residual network to construct the mapping relationship between spatial positioning cues and the azimuth of the sound source. This positioning model can accelerate the features in the circulation network, reduce feature loss, and lower the training difficulty; the present invention can complete the training process of the positioning model CRN network offline, save the trained network in memory, and only need one frame of signal to achieve real-time sound source positioning during testing. Compared with the traditional SRP-PHAT algorithm and the positioning algorithm based on deep neural network, the algorithm of the present invention significantly improves the positioning performance in complex acoustic environments and has good generalization ability for the spatial structure of the sound source, reverberation, and noise. Description of the Drawings
[0069] Figure 1 is a flowchart of a method for sound source localization of a microphone array provided by an embodiment of the present invention;
[0070] Figure 2 is a block diagram of the CRN model structure provided by an embodiment of the present invention;
[0071] Figure 3 is a block diagram of the residual block structure provided by an embodiment of the present invention;
[0072] Figure 4 、 Figure 5 is a positioning effect diagram of various algorithms when the test environment and the training environment are the same provided by an embodiment of the present invention;
[0073] Figure 6 、 Figure 7 is a positioning result diagram in a non-training noise environment provided by an embodiment of the present invention;
[0074] Figure 8 、 Figure 9 is a positioning result diagram in a non-training reverberation environment provided by an embodiment of the present invention. Detailed Embodiments
[0075] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0076] Embodiment 1
[0077] This embodiment introduces a method for sound source localization of a microphone array, including:
[0078] Obtain a test signal;
[0079] Preprocess the test signal to obtain a single-frame test signal;
[0080] Extract the spatial localization cues of the single-frame test signal and use them as test samples;
[0081] Input the test samples into the pre-constructed and trained CRN model for testing, and obtain the probability that the test signal belongs to each azimuth angle. Among them, the azimuth angle with the maximum probability is taken as the azimuth angle estimation value of this frame of signal.
[0082] The method for localizing the sound source of the microphone array provided in this embodiment specifically involves the following steps in its application process:
[0083] Step 1: Convolve the clean speech signal with the room impulse responses at different azimuth angles, and add different degrees of noise and reverberation to generate multiple directional speech signals at different azimuths:
[0084] x m (t) = h m (t) * s(t) + v m (t), m = 1, 2,..., M
[0085] Where x m (t) represents the speech signal received by the m-th microphone at a specified azimuth. m is the serial number of the microphone element, m = 1, 2,..., M, and M is the number of microphone elements. s(t) is the clean speech, and h m (t) represents the room impulse response from the specified sound source azimuth to the m-th microphone. h m (t) is related to the sound source azimuth and room reverberation, and v m (t) represents the noise.
[0086] In this embodiment, it is set that the microphone array is a uniform circular array composed of 6 omnidirectional microphones, and the array radius is 0.1 m. It is set that the sound source and the microphone array are in the same horizontal plane, the sound source is in the far field of the array, the due front of the horizontal plane is defined as 90°, the range of the sound source azimuth angle is [0°, 360°), and the interval is 10°. Then the number of training azimuths is 36. The reverberation time of the training data includes 0.5 s and 0.8 s, and the room impulse responses h m (t) at different azimuth angles under different reverberation times are generated by the Image algorithm. v m (t) is Gaussian white noise, and the signal-to-noise ratios of the training data include 0 dB, 5 dB, 10 dB, 15 dB, and 20 dB.
[0087] Step 2: Preprocess the microphone array signals obtained in Step 1 to obtain single-frame signals.
[0088] The preprocessing includes frame segmentation and windowing, where:
[0089] The frame segmentation method is: using a preset frame length and frame shift, segment the time-domain signal x of the m-th element m(t) is divided into multiple single-frame signals x m (iN + n), where i is the frame sequence number, n represents the sampling sequence number within a frame, 0 ≤ n < N, and N is the frame length. In this embodiment, the sampling rate f of the voice signal s is 16 kHz, the frame length N taken is 512 (32 ms), and the frame shift is 0.
[0090] The windowing method is: x m (i, n) = w H (n)x m (iN + n)
[0091] where x m (i, n) is the signal of the i-th frame of the m-th array element after windowing processing,
[0092] and it is a Hamming window.
[0093] Step 3: Extract the spatial positioning clues of the array signal. Specifically, it includes:
[0094] (3 - 1) Perform discrete Fourier transform on each single-frame signal obtained in Step 2 to convert the time-domain signal to the frequency-domain signal.
[0095] The calculation formula of the discrete Fourier transform is:
[0096]
[0097] where X m (i, k) is the discrete Fourier transform of x m (i, n), representing the frequency-domain signal of the i-th frame of the m-th array element, k is the frequency point, K is the length of the discrete Fourier transform, K = 2N, and DFT(·) represents the discrete Fourier transform. In this embodiment, the Fourier transform length is set to 1024.
[0098] (3 - 2) Design a Gammatone filter bank. g j (t) is the impulse response function of the j-th Gammatone filter, and its expression is
[0099]
[0100] where j represents the filter sequence number; C is the filter gain; t represents continuous time; a is the filter order; represents the phase; f j represents the center frequency of the j-th filter; b j represents the filter attenuation factor, b j The calculation formula is:
[0101] b j= 1.109ERB(f j )
[0102] ERB(f j ) = 24.7(4.37f j / 1000 + 1)
[0103] In this embodiment, the order a is 4, the phase is set to 0, the number of sub-band filters is 32, that is, j = 1, 2, ···, 32, and the center frequency f j of the filter ranges from [200 Hz, 8000 Hz]
[0104] Perform discrete Fourier transform on each Gammatone filter to obtain its frequency-domain expression:
[0105]
[0106] (3 - 3) Calculate the sub-band generalized cross-correlation function of each frame of signal, and its calculation formula is as follows:
[0107]
[0108] where R mn (i, j, τ) represents the generalized cross-correlation function of the m-th array element and the n-th array element in the i-th frame and the j-th sub-band.
[0109] (3 - 4) Obtain the sub-band time-delay difference of each frame of signal, and its expression is as follows:
[0110]
[0111] where T mn (i, j) represents the time-delay difference of the m-th array element and the n-th array element in the i-th frame and the j-th sub-band.
[0112] (3 - 5) Calculate the sub-band SRP-PHAT function of each frame of signal, and the calculation formula is as follows
[0113]
[0114] where P(i, j, r) represents the SRP-PHAT power value of the i-th frame of signal and the j-th sub-band when the beam direction of the array is r; τ mn (r) represents the time difference between the sound wave propagating from the beam direction r to the m-th microphone and the n-th microphone, and its calculation formula is:
[0115]
[0116] where r represents the coordinate of the beam direction, r mDenote the position coordinates of the m-th microphone, c is the speed of sound in air, approximately 342 m / s at room temperature, and f s is the signal sampling rate.
[0117] In this embodiment, it is set that the sound source and the microphone array are in the same horizontal plane, and the sound source is in the far field of the array, then τ mn (r)'s equivalent calculation formula is:
[0118]
[0119] where ξ = [cosθ, sinθ] T , and θ is the azimuth angle of the beam direction r. τ mn (r) has nothing to do with the received signal, so it can be calculated offline and saved in memory.
[0120] Normalize the sub-band SRP-PHAT function, and the calculation formula is as follows:
[0121]
[0122] (3-6) Combine the time delay differences and SRP-PHAT functions of all sub-bands within the same frame to form a feature matrix, and obtain the spatial clue of the mixed feature. The expression is as follows:
[0123]
[0124] where y train (i) represents the spatial localization clue of the i-th frame signal, J is the number of sub-bands, and J = 32 in this embodiment. In this embodiment, the azimuth range of the beam pointing is [0°, 360°), and the interval is 10°, so L = 36. In this embodiment, the number of microphones M is taken as 6, so there are a total of pairs of microphones. Therefore, the dimension of the localization clue for each sub-band in each frame is 15 + 36 = 51, and the dimension of the feature matrix is 51×32.
[0125] Step Four: Prepare the training set: According to Steps One to Three, extract the spatial feature parameters of the directional speech data in all training environments (for the implementation settings of the training environment, see Step One in detail), use them as the training samples of the CRN, and at the same time mark the corresponding azimuth of each sample as the class label of the sample.
[0126] Step Five: Construct a CRN model, use the training samples and class labels obtained in Step Four as the training data set of the CNN, and perform training to obtain the CNN model. Specifically, it includes:
[0127] (5-1) Set the CNN model structure.
[0128] The CRN model structure framework adopted in the present invention is as Figure 2As shown in the figure, it includes an input layer, two residual blocks, a pooling layer, two fully connected layers, and a final output layer. Among them, the structure of each residual block is as shown in Figure 3 shown. The residual block is composed of multiple convolutional layers and batch normalization (BN) layers. Each residual block structure contains two BN layers and two convolutional layers, and processes the input in the order of BN first, followed by ReLU, and finally the convolutional layer.
[0129] The input signal of the input layer is the two-dimensional sub-band. In this embodiment, J = 32, L = 36, M = 6. The size of the convolutional kernel in the convolutional layer of the residual block in the CRN model is 3×3, the stride is 1×1, and zero-padding is used to keep the dimension of the feature parameters unchanged before and after convolution. The number of hidden units in the first fully connected layer is 128, and the number of hidden units in the second fully connected layer is 36, which is also the number of output azimuth angles. The output layer uses Softmax, and the loss function is the cross-entropy loss function.
[0130] (5-2) Train the network parameters of the CRN model.
[0131] In the present invention, the Adam optimizer is used to continuously reduce the loss function during the training process of the model. During the CRN training process, information propagates forward, errors propagate backward, and the model parameters are updated accordingly. The present invention uses Xavier to initialize the model parameters. Assuming that the sizes of the input weights and output weights of the weight layer W are n j and n j+1 respectively, the formula for Xavier initialization is as follows:
[0132]
[0133] where U represents a uniform distribution, and the symbol ∼ means that the weight layer W follows the uniform distribution of U.
[0134] During training, the initial learning rate is set to 0.001, the batch size of the data is set to 200, the value of ε in the BN layer is 0.001, and the decay coefficient is taken as 0.999. The present invention uses the method of cross-validation. In each round of iteration, the training data is randomly divided into two parts: 70% training set and 30% validation set. Through repeated cross-validation, the cross-entropy loss function is used to measure the quality of the model, and finally the training stage of the model is completed.
[0135] Step six: Process the test signal according to steps two and three to obtain the spatial localization clue y test (i) of the single-frame test signal, and use it as a test sample.
[0136] Step 7: Use the test sample as the input feature of the CRN model trained in Step 5. The CRN outputs the probability that the test signal belongs to each azimuth angle, and take the azimuth with the maximum probability as the azimuth angle estimation value of the signal of this frame.
[0137] The present invention uses hybrid features (sub-band time delay difference and sub-band SRP-PHAT spatial spectrum) as spatial positioning clues, and these feature clues have strong robustness and spatial information representation ability. The present invention uses a convolutional residual network to construct the mapping relationship between the spatial positioning clues and the sound source azimuth. This positioning model can accelerate the features in the circulation network, reduce feature loss, and reduce the training difficulty. The present invention can complete the training process of the CRN network of the positioning model offline, save the trained network in the memory, and only one frame of signal is required for real-time sound source positioning during testing. Compared with the traditional SRP-PHAT algorithm and the positioning algorithm based on a deep neural network, the algorithm of the present invention significantly improves the positioning performance in a complex acoustic environment, and has good generalization ability for the spatial structure, reverberation and noise of the sound source.
[0138] Figure 4 、 Figure 5 shows the positioning effects of various algorithms when the test environment is consistent with the training environment. It can be seen from the figure that the positioning success rate of the algorithm of the present invention is higher than that of the traditional SRP-PHAT and the positioning algorithm based on a deep neural network. Figure 6 、 Figure 7 、 Figure 8 and Figure 9 show the positioning effects of various algorithms when the test environment is inconsistent with the training environment. Figure 6 、 Figure 7 are the positioning results in a non-training noise environment, Figure 8 、 Figure 9 are the positioning results in a non-training reverberation environment. It can be seen from the figure that even in a non-training environment, the success rate of the algorithm of the present invention is still higher than that of the traditional SRP-PHAT algorithm and the positioning algorithm based on a deep neural network, indicating that the method of the present invention has better robustness and generalization ability for unknown environments.
[0139] Embodiment 2
[0140] The present embodiment provides a microphone array sound source positioning device, including:
[0141] An acquisition unit, configured to acquire a test signal;
[0142] A preprocessing unit, configured to preprocess the test signal to obtain a single-frame test signal;
[0143] An extraction unit, configured to extract the spatial positioning clue of the single-frame test signal and use it as a test sample;
[0144] A test unit for inputting the test sample into a pre-constructed and trained CRN model for testing, obtaining the probability that the test signal belongs to each azimuth angle, and taking the azimuth with the maximum probability as the azimuth angle estimation value of the frame signal.
[0145] Further, the test unit includes a module for constructing and training a CRN model, and the module for constructing and training a CRN model includes:
[0146] A microphone array signal generation module for convolving a pure speech signal with room impulse responses at different azimuth angles, adding different degrees of noise and reverberation, and generating multiple microphone array signals;
[0147] A preprocessing module for preprocessing the multiple microphone array signals to obtain multiple single-frame signals;
[0148] An extraction module for extracting spatial localization clues of multiple single-frame signals, using them as training samples of the CRN model, and simultaneously marking the corresponding azimuth of each sample as the class label of the sample;
[0149] A construction and training module for constructing a CRN model and training the training samples and class labels as the training data set of the CRN model.
[0150] Among them, the CRN model includes an input layer, two residual blocks, a pooling layer, two fully connected layers, and a final output layer. Among them, the residual block is composed of multiple convolutional layers and batch normalization layers. Each residual block structure contains two batch normalization layers and two convolutional layers, and processes the input in the order of first batch normalization layer, then ReLU, and finally convolutional layer. Among them, the output layer uses Softmax, and the loss function is the cross-entropy function.
[0151] Embodiment 3
[0152] This embodiment provides a microphone array sound source localization device, including a processor and a storage medium;
[0153] The storage medium is used to store instructions;
[0154] The processor is used to operate according to the instructions to execute the steps of the method according to any one of Embodiment 1.
[0155] Embodiment 4
[0156] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of Embodiment 1.
[0157] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for localizing the sound source of a microphone array, characterized in that, it includes: Obtain a test signal; Preprocess the test signal to obtain a single-frame test signal; Extract the spatial localization clues of the single-frame test signal and use it as a test sample; Input the test sample into a pre-constructed and trained Convolutional Residual Network (CRN) model for testing, and obtain the probability that the test signal belongs to each azimuth angle. Among them, the azimuth angle with the maximum probability is taken as the azimuth angle estimation value of this frame of signal; The method for constructing and training the Convolutional Residual Network (CRN) model includes: Convolve the clean speech signal with the room impulse response at different azimuth angles, and add different degrees of noise and reverberation to generate multiple microphone array signals; Preprocess the multiple microphone array signals to obtain multiple single-frame signals; Extract the spatial localization clues of the multiple single-frame signals, use it as the training sample of the Convolutional Residual Network (CRN) model, and at the same time mark the corresponding azimuth of each sample as the class label of the sample; Construct a Convolutional Residual Network (CRN) model, and use the training sample and class label as the training data set of the Convolutional Residual Network (CRN) model for training; The extraction of the spatial localization clues of the multiple single-frame signals includes: Perform a discrete Fourier transform on each single-frame signal to convert the time-domain signal to a frequency-domain signal; The calculation formula of the discrete Fourier transform is: Where X m (i,k) is x m (i,n) is the discrete Fourier transform of, representing the frequency-domain signal of the m-th array element in the i-th frame, k is the frequency point, x m (i,n) is the signal of the m-th array element in the i-th frame after windowing, K is the length of the discrete Fourier transform, K = 2N, N is the frame length, and DFT(·) represents the discrete Fourier transform; Design the Gammatone filter bank, g j (t) is the impulse response function of the j-th Gammatone filter, and its expression is: Among them, j represents the serial number of the filter; C is the filter gain; t represents continuous time; a is the order of the filter; represents the phase; f j represents the center frequency of the j-th filter; b j represents the filter attenuation factor, b j The calculation formula is: b j = 1.109ERB(f j ) ERB(f j ) = 24.7(4.37f j / 1000 + 1) Perform a discrete Fourier transform on each Gammatone filter to obtain its frequency-domain expression: Calculate the sub-band generalized cross-correlation function of each frame of signal, and its calculation formula is as follows: Among them, R mn (i, j, τ) represents the generalized cross-correlation function between the m-th array element and the n-th array element in the i-th frame and the j-th sub-band; Obtain the sub-band time delay difference of each frame of signal, and its expression is as follows: Among them, Τ mn (i,j) represents the time delay difference between the m-th array element and the n-th array element in the i-th frame and the j-th sub-band; Calculate the sub-band SRP-PHAT function of each frame of signal, and the calculation formula is as follows: Among them, P(i, j, r) represents the SRP-PHAT power value of the j-th sub-band of the i-th frame signal when the beam direction of the array is r; τ mn (r) represents the time difference between the sound wave propagating from the beam direction r to the m-th microphone and the n-th microphone, and its calculation formula is: where r represents the coordinates of the beam direction, r m represents the position coordinates of the m-th microphone, c is the speed of sound in air, f s is the signal sampling rate; Set the sound source and the microphone array to be on the same horizontal plane, with the sound source in the far field of the array, then the equivalent calculation formula of τ mn (r) is: where ξ = [cosθ, sinθ] T , θ is the azimuth angle of the beam direction r, τ mn (r) has nothing to do with the received signal and can therefore be calculated offline and saved in memory; Perform a normalization process on the sub-band SRP-PHAT function, and the calculation formula is as follows: Form a feature matrix with the time delay difference and SRP-PHAT function of all sub-bands within the same frame to obtain the spatial clues of the mixed features, and the expression is as follows: where y train (i) represents the spatial localization clue of the i-th frame signal, and J is the number of subbands.
2. The method for localizing the sound source of a microphone array according to claim 1, characterized in that, The convolution of the clean speech signal with the room impulse response at different azimuth angles, and adding different degrees of noise and reverberation to generate multiple microphone array signals, the formula is as follows: x m y(t) = h m (t) * s(t) + v m (t), m = 1, 2,..., M where x m (t) represents the speech signal in a specified direction received by the m-th microphone, where m is the serial number of the microphone element, m = 1, 2, …, M, M is the number of microphone elements, s(t) is the clean speech, and h m (t) represents the room impulse response from the specified sound source direction to the m-th microphone. h m (t) is related to the sound source direction and room reverberation, and v m (t) represents noise.
3. The method for localizing the sound source of a microphone array according to claim 1, characterized in that, The preprocessing of the multiple microphone array signals to obtain multiple single-frame signals includes: The preprocessing includes framing and windowing, where: The framing method is as follows: Using a preset framing length and frame shift, the time-domain signal x of the m-th array element m (t) is divided into multiple single-frame signals x m (iN + n), where i is the frame sequence number, n represents the sampling sequence number within a frame, 0 ≤ n < N, and N is the frame length; The windowing method is: x m (i,n) = w H (n)x m (iN + n) where x m (i,n) is the signal of the m-th array element in the i-th frame after windowing processing is a Hamming window.
4. The method for localizing the sound source of a microphone array according to claim 3, characterized in that: The method for preprocessing the test signal to obtain a single-frame test signal is the same as the method for preprocessing the multiple microphone array signals to obtain multiple single-frame signals; The method for extracting the spatial localization clues of the single-frame test signal is the same as the method for extracting the spatial localization clues of the multiple single-frame signals.
5. The method for localizing the sound source of a microphone array according to claim 1, characterized in that: The convolutional residual network CRN model includes an input layer, two residual blocks, a pooling layer, two fully connected layers, and a final output layer. Among them, each residual block is composed of multiple convolutional layers and batch normalization layers. Each residual block structure contains two batch normalization layers and two convolutional layers, and processes the input in the order of first batch normalization layer, then ReLU, and finally convolutional layer. Among them, the output layer uses Softmax, and the loss function is the cross-entropy function.
6. A microphone array sound source localization device, which adopts the microphone array sound source localization method described in claim 1. It is characterized in that: It includes: An acquisition unit for acquiring a test signal; A preprocessing unit for preprocessing the test signal to obtain a single-frame test signal; An extraction unit for extracting the spatial localization clue of the single-frame test signal and using it as a test sample; A test unit for inputting the test sample into a pre-constructed and trained convolutional residual network CRN model for testing, and obtaining the probability that the test signal belongs to each azimuth angle. Among them, the azimuth angle with the largest probability is taken as the azimuth angle estimation value of this frame of signal.
7. A microphone array sound source localization device It is characterized in that: It includes a processor and a storage medium; The storage medium is used for storing instructions; The processor is used to operate according to the instructions to execute the steps of the method described in any one of claims 1 to 5.
8. A computer-readable storage medium, on which a computer program is stored. It is characterized in that: When the program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Sound source localization method based on convolutional neural network and sub-band SRP-PHAT spatial spectrum
CN112904279A