Multi-sound-source arrival direction estimation method and device based on frequency focusing spatial spectrum
By using the multi-sound source arrival direction estimation method based on frequency-focused spatial spectrum in multi-sound sources and complex acoustic scenarios, the influence of noise and reverb on DOA estimation performance is solved, and a high-precision and robust multi-sound source positioning and separation effect is achieved.
Patent Information
- Application Number
- CN202411976995.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
Noise and reverb have a serious impact on the multi-source direction of arrival (DOA) estimation performance of microphone arrays in multi-sound source and complex acoustic scenarios, and existing methods exhibit insufficient accuracy and robustness in these environments.
Using a multi-sound source arrival direction estimation method based on frequency focusing spatial spectrum, a microphone array data set is generated through a preset simulation environment, a training data set is constructed and sound source masking training is performed, and a mask estimation network model is generated. Then, the multi-sound source signal is sub-bandwidth partitioned through frequency band superposition, the weighted focus covariance matrix is obtained, the broadband frequency focus spatial spectrum characteristics are constructed, and the multi-sound source DOA spatial spectrum is estimated using the convolutional neural network model.
It improves the DOA estimation accuracy and robustness in a multi-sound source environment with noisy and reverb, reduces the influence of nonlinear distortion in neural networks, and enhances the stability and fault tolerance of the algorithm.
Smart Images

Figure CN119936786A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of acoustic signal processing and sound source localization, and in particular to a method and device for estimating the arrival directions of multiple sound sources based on frequency-focused spatial spectrum. Background Art
[0002] In recent years, online audio-visual conference scenarios based on microphone arrays have become a hot topic in acoustic research, especially in the field of multi-source DOA (Direction of arrival) estimation. This technology provides support for subsequent tasks such as speech enhancement, multi-speaker speech separation, spatial audio coding, and sound field reconstruction by obtaining the spatial information of the speaker, thereby improving the robustness of speech communication, recognition, and conference content transcription. However, in indoor conference environments, factors such as noise and reverberation have a serious impact on the performance of DOA estimation, especially in multi-source and complex acoustic scenes, where the random distribution of sound sources and signal overlap make DOA estimation more difficult. Although there are some signal modeling-based methods, such as SRP, MUSIC, and PHAT, which can stably perform DOA estimation under ideal conditions, their performance drops sharply under noise and reverberation interference. Therefore, improving the accuracy and robustness of DOA estimation in complex environments has become a difficult point in current research.
[0003] In order to solve the impact of reverberation on DOA estimation performance, neural networks have been introduced into this field in recent years, and some DOA estimation methods based on supervised learning have been proposed, mainly including three categories: first, using neural networks to correct pre-extracted features, such as enhancing inter-channel phase features to improve estimation accuracy; second, classifying extracted feature labels through neural networks to improve robustness in reverberant environments, such as CNN-based SPR-PHAT and MUSIC algorithms; third, based on the sparsity of talker time-frequency units, using neural networks for signal separation to reduce the impact of noise and reverberation. Although these neural network-based DOA estimation methods have shown relatively superior performance, they still face some problems, such as the poor generalization ability of network models in indoor environments, the possible introduction of nonlinear distortion by masking enhancement methods, and poor robustness in multi-source scenarios.
[0004] Therefore, one or more methods are needed to solve the above problems.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The present disclosure shows a method and device for estimating the arrival direction of multiple sound sources based on frequency-focused spatial spectrum.
[0007] In a first aspect, the present disclosure provides a method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum, the method comprising:
[0008] Using a preset simulation environment and a microphone array signal simulator, a microphone array data set is generated from a single sound source clean speech corpus, and a training data set is constructed based on the microphone array data set;
[0009] By performing sound source masking training on the training data set, a mask estimation network model is generated; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking;
[0010] According to the mask estimation network model, an enhanced multi-sound source signal is obtained, and the enhanced multi-sound source signal is divided into sub-bandwidths by frequency band overlapping to obtain a weighted focusing covariance matrix of each sub-bandwidth;
[0011] Based on the weighted focusing covariance matrix of each sub-bandwidth, a sub-bandwidth frequency focusing covariance matrix obtained by beamforming is used to construct a broadband frequency focusing spatial spectrum feature, and a convolutional neural network model is constructed based on the broadband frequency focusing spatial spectrum feature, and the convolutional neural network model is used to estimate the DOA spatial spectrum of multiple sound sources;
[0012] The mask estimation network model and the multi-sound source DOA spatial spectrum estimation network model are used to decode and obtain the multi-sound source DOA spatial spectrum, and significant peaks are extracted to complete the estimation of the multi-sound source DOA.
[0013] In an exemplary embodiment of the present disclosure, the microphone array data set includes: a clean microphone array speech signal, a noisy and reverberant microphone array signal, and a sound source DOA label.
[0014] In an exemplary embodiment of the present disclosure, constructing a training data set based on the microphone array data set includes:
[0015] Using the simulation environment, constructing masking training data by performing short-time Fourier transform on the clean microphone array speech signal and the noisy and reverberant microphone array signal;
[0016] According to the short-time Fourier transform structure of the noisy and reverberant microphone array signal, the activity area of the clean microphone array speech is detected and labeled by an endpoint check algorithm to construct sound source DOA spatial spectrum training data;
[0017] A training data set is generated by integrating the masking training data and the sound source DOA spatial spectrum training data.
[0018] In an exemplary embodiment of the present disclosure, the enhanced multi-sound source signal is divided into sub-bandwidths by means of frequency band overlap, including:
[0019] The broadband spatial spectrum is obtained by focusing the covariance matrix of the sub-bandwidth as an intermediate representation of the spatial orientation information of the sound source;
[0020] The signal bandwidth is divided into multiple sub-bandwidths in a sub-bandwidth overlapping manner, and the speech existence probability of each sub-bandwidth is obtained based on a speech existence probability algorithm, and the signals of each sub-bandwidth are weighted to obtain a sub-bandwidth weighted focusing covariance matrix.
[0021] In an exemplary embodiment of the present disclosure, a broadband frequency-focused spatial spectrum feature is constructed by using a sub-bandwidth frequency-focused covariance matrix obtained by beamforming, including:
[0022] The spatial spectrum of the sub-bandwidth is modeled based on the minimum variance distortion-free response beamforming method, and the sub-bandwidth frequency-focused spatial spectrum characteristics based on beamforming are obtained.
[0023] In an exemplary embodiment of the present disclosure, the convolutional neural network includes:
[0024] The convolutional neural network includes a convolutional layer for extracting identification features and a fully connected layer for mapping the DOA spatial spectrum of the sound source wave.
[0025] In an exemplary embodiment of the present disclosure, the mask estimation network model and the multi-sound source DOA spatial spectrum estimation network model are used to decode and obtain the multi-sound source DOA spatial spectrum, including:
[0026] Based on the decoding process of the mask estimation network model and the sound source DOA estimation network model, the DOA prediction result of the multi-sound source wave is obtained by calculating the significant peak of the multi-sound source DOA spatial spectrum;
[0027] Based on the sound source wave DOA prediction result, the DOA of multiple sound sources is estimated.
[0028] In a second aspect, the present disclosure provides a device for estimating directions of arrival of multiple sound sources based on frequency-focused spatial spectrum, the device comprising:
[0029] A data construction module, used to generate a microphone array data set from a single sound source clean speech corpus using a preset simulation environment and a microphone array signal simulator, and to construct a training data set based on the microphone array data set;
[0030] A model building module, used to generate a mask estimation network model by performing sound source masking training on the training data set; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking;
[0031] A covariance matrix acquisition module, used to acquire the enhanced multi-sound source signal according to the mask estimation network model, and divide the enhanced multi-sound source signal into sub-bandwidths by frequency band overlap to acquire a weighted focusing covariance matrix of each sub-bandwidth;
[0032] A spatial spectrum training module, for constructing broadband frequency-focused spatial spectrum features based on the weighted focusing covariance matrix of each sub-bandwidth and the sub-bandwidth frequency-focused covariance matrix obtained by beamforming, and constructing a convolutional neural network model based on the broadband frequency-focused spatial spectrum features, and estimating the DOA spatial spectrum of multiple sound sources using the convolutional neural network model;
[0033] The spatial spectrum inference module is used to predict the sound source DOA spatial spectrum according to the decoding process of the mask estimation network model and the sound source DOA estimation network model, and complete the estimation of multiple sound source DOAs.
[0034] In a third aspect, the present disclosure shows an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.
[0035] In a fourth aspect, the present disclosure shows a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.
[0036] In a fifth aspect, the present disclosure shows a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any of the above aspects.
[0037] The method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum disclosed in the present invention includes: first, using a preset simulation environment and a microphone array signal simulator to generate a microphone array data set and a training data set; second, constructing a mask estimation network model through sound source masking training to estimate the sound mask information; then, dividing the multi-sound source signal into sub-bandwidths by frequency band overlap to obtain a weighted focusing covariance matrix; then, obtaining bandwidth-focused spatial spectrum features based on beamforming and sub-bandwidth focusing covariance matrix, and training a convolutional neural network for estimating the sound source DOA based on this feature. Finally, the sound source DOA spatial spectrum is predicted according to the decoding process of the network model to complete the estimation of the DOA of multiple sound sources. The present invention is based on the intermediate representation of the broadband frequency-focused spatial spectrum, which reduces the influence of the nonlinear distortion of the neural network on the DOA estimation performance; the optimized broadband frequency-focused spatial spectrum is used to map the DOA spatial spectrum, which improves the stability of the algorithm performance and the fault tolerance of the DOA estimation; the proposed algorithm has strong robustness in a multi-sound source environment containing noise and reverberation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The present invention discloses a method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum.
[0039] Figure 2 This is a flow chart of a model building method, training method, testing and reasoning method for a method for estimating the arrival direction of multiple sound sources based on a frequency-focused spatial spectrum disclosed in the present invention.
[0040] Figure 3 It is a schematic diagram of a simulation environment setting of a microphone array signal of a method for estimating the arrival direction of multiple sound sources based on frequency-focused spatial spectrum disclosed in the present invention.
[0041] Figure 4 This is a network structure diagram of mask training of a multi-sound source arrival direction estimation method based on frequency-focused spatial spectrum and a network structure diagram of sound source wave DOA spatial spectrum training.
[0042] Figure 5 It is a structural block diagram of a device for estimating the arrival directions of multiple sound sources based on frequency-focused spatial spectrum disclosed in the present invention.
[0043] Figure 6 is a block diagram of an electronic device disclosed herein.
[0044] Figure 7 is a block diagram of a computer-readable storage medium of the present disclosure. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0046]
Term explanation
[0047] Masking in acoustics refers to the process of processing signals using an acoustic mask, which is often expressed as a ratio of the target signal to the observed signal. Therefore, it is mainly used in signal enhancement, signal separation, target signal extraction and other processing.
[0048] Frequency focusing is a processing method that moves signals of different frequency bands to the same frequency band through a focusing matrix in the frequency domain. It is widely used in signal processing fields such as radar, sonar, antenna and voice.
[0049] Convolutional Neural Networks (CNN) is a type of feedforward neural network that includes convolution calculations and has a deep structure. It can perform translation-invariant classification of input information according to a hierarchical structure. It is one of the most representative neural networks and has a wide range of applications in the field of image processing.
[0050] Reference Figure 1 , shows a flowchart of a method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum of the present disclosure, which can be applied to electronic devices, wherein the method specifically may include the following steps:
[0051] Step S110, using a preset simulation environment and a microphone array signal simulator, generating a microphone array data set from a single sound source clean speech corpus, and constructing a training data set based on the microphone array data set;
[0052] Step S120, generating a mask estimation network model by performing sound source masking training on the training data set; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking;
[0053] Step S130, obtaining an enhanced multi-sound source signal according to the mask estimation network model, and dividing the enhanced multi-sound source signal into sub-bandwidths by means of frequency band overlap to obtain a weighted focusing covariance matrix of each sub-bandwidth;
[0054] Step S140, based on the weighted focusing covariance matrix of each sub-bandwidth, a sub-bandwidth focusing covariance matrix obtained by beamforming is used to construct a broadband frequency focusing spatial spectrum feature, and a convolutional neural network model is constructed based on the broadband frequency focusing spatial spectrum feature, and the convolutional neural network model is used to estimate the multi-sound source DOA spatial spectrum;
[0055] Step S150, using the mask estimation network model and the multi-sound source DOA spatial spectrum estimation network model, decoding to obtain the multi-sound source DOA spatial spectrum, extracting significant peaks to complete the estimation of the multi-sound source DOA.
[0056] The method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum provided by the present invention is based on the intermediate representation of broadband frequency-focused spatial spectrum, which reduces the influence of nonlinear distortion of neural network on DOA estimation performance; uses optimized broadband frequency-focused spatial spectrum to map DOA spatial spectrum, which improves the stability of algorithm performance and the fault tolerance of DOA estimation; the proposed algorithm has strong robustness in a noisy and reverberant multi-sound source environment.
[0057] In the embodiment of this example, in order to reduce the sensitivity of DOA estimation to nonlinear distortion caused by masking enhancement, the present disclosure adopts the broadband frequency-focused spatial spectrum features of masked enhanced speech as the intermediate representation of the spatial orientation information of the sound source, rather than directly using the masked enhanced speech to estimate the DOA information of the sound source; in order to improve the representation power of the broadband frequency-focused spatial spectrum for the spatial orientation information of the sound source, a series of optimization processes are performed, such as improving the frequency resolution of the broadband spatial spectrum features by dividing the sub-bandwidth, using the mean of the snapshot inner mask to weight the focused covariance matrix to reduce the influence of the noise-dominated frequency band, and using the beamforming method to construct the spatial spectrum to avoid the problem of estimating the number of sound sources; in order to improve the spatial resolution and fault tolerance of the DOA estimation algorithm, the clean speech and endpoint detection results are used to construct the DOA spatial spectrum label at a specified resolution, and the broadband spatial spectrum features are used to map the narrowband DOA spatial spectrum label through a convolutional neural network to achieve robust DOA estimation performance in complex environments.
[0058] In this exemplary embodiment, if Figure 2 As shown, the present disclosure includes the following steps:
[0059] Step S110: Generate data based on Pyroomacoustics and prepare a training data set.
[0060] 1) The simulation environment for microphone array signal generation is as follows Figure 3 As shown in the figure, the length of the room is randomly selected between 5m and 12m; the width of the room is randomly selected between 3m and 8m, and the width is required to be ≥ length / 2; the height of the room is randomly selected between 2.8m and 5m. The number of elements of the circular microphone array is 6, and the radius is 0.036m. The center of the array is placed at a random position in the smaller dotted box in the center of the room. The center of the dotted box is aligned with the center of the room. The length and width are both 0.4m. The dark gray larger dotted box is the conference table placement area. The length of the conference table is between 2.5m and 8m and is constrained to be greater than or equal to 1 / 2 of the room length. The width of the conference table is between 1m and 4m and is constrained to be greater than or equal to 1 / 3 of the room width. The height of the conference table is randomly selected between 1m and 1.5m. The light gray area is the speaker activity area, which is greater than 0.25m away from the wall. At the same time, the number of speakers is between 1 and 4, and the speaker position is randomly generated in the gray area, with a minimum angle of 5° between speakers compared to the center of the microphone array.
[0061] 2) Generate microphone array data set. Use the clean speech data set and microphone array signal generator to generate the corresponding clean microphone array speech, noisy reverberation microphone array signal, and sound source DOA label; among which, the signal-to-noise ratio (SNR) is randomly generated between 0dB and 20dB, the reverberation time RT60 is randomly generated between 0ms and 1200ms, and the signal-to-interference ratio (SIR) between sound sources is randomly generated between –5dB and 5dB.
[0062] 3) Construction of masking training data. Perform short-time Fourier transform (STFT) on the generated clean speech, noisy and reverberant speech of each channel to generate the corresponding input complex spectrum and the corresponding sound source complex ideal ratio masking (cIRM). The definition of cIRM and the network structure used for training are given in step S120.
[0063] 4) Construction of DOA spatial spectrum training data. First, based on the STFT structure of the noisy and reverberant speech of each channel, the broadband frequency-focused spatial spectrum features are calculated and used as the input of DOA spatial spectrum estimation. For detailed calculation steps, see steps S130 to S150; secondly, the endpoint detection (VAD) algorithm is used to detect the active area of the clean speech of each sound source, and the DOA spatial pseudo spectrum (SPS) label is generated for the speech active area by formula (1), which is used as the ideal value of DOA spatial spectrum estimation, so as to calculate the loss value and optimize the network parameters.
[0064]
[0065] in, is the constructed DOA spatial pseudo-spectrum; is the angle of spatial spectrum traversal. Different angle traversal intervals can be set to obtain DOA spatial pseudo-spectra with different resolutions. The traversal interval is generally set to 5°, that is, when the DOA is correctly resolved, the maximum error of the DOA estimation is 2.5°; q and Q are the sound source index and the number of sound sources emitting simultaneously, respectively, θ q is the actual DOA angle of the qth sound source.
[0066] Step S120: construct a network model for sound source masking training.
[0067] A large number of recent studies have shown that the multi-source DOA estimation problem can be made easier by separating individual sound sources one by one through masking. However, the number of sound sources is often unknown, and training multiple maskings makes network training more difficult. Therefore, the present disclosure only trains multi-source clean speech signals (i.e., direct sound signals), that is, the cIRM generated in step S110 is shown in equations (2) to (3):
[0068]
[0069] Wherein, the subscript m represents the index of the number of microphone array channels; the subscript q = 1, 2, ..., Q, q is the index of the number of sound sources, Q is the total number of sound sources; k = 1, 2, ..., K, l = 1, 2, ..., L, k and l are the indexes of frequency and frame respectively, K and L are the number of frequency and frame respectively; X m (k, l) is the mixed signal of Q sound sources collected by the mth channel, S q (k,l) is the original signal of the qth sound source, H m,q (0) is the direct sound transfer function from the qth sound source to the mth microphone. The symbols "Ι" and "Ρ" represent the imaginary and real part operations, respectively, and i represents the imaginary unit. It can be seen from formula (2) that the constructed cIRM is the ratio of the mixed signal of the direct sound of Q sound sources to the observed signal, that is, its purpose is to remove noise and reverberation and retain the direct sound signals of Q sounds.
[0070] Based on the mask of formula (2), the constructed mask estimation network model is as follows: Figure 4 As shown. Figure 4 In , the input signal is the complete set features of the past two frames, the current frame, and the next two frames, with a dimension of (246*5)×1; three hidden layers are used to map complex ratio masking, and the number of neurons in each hidden layer is 512; the output layer dimension is 257×2, that is, the real part of the mask is 257×1 and the imaginary part is 257×1. Based on formula (2), the masked enhanced signal Z of each channel can be obtained by formula (4): m (k,l):
[0071] Z m (k,l)=λ m (k,l)Y m (k,l) (4)
[0072] The mean square error of cIRM is used as the loss function Λ, as shown in Equation (5).
[0073]
[0074] in, To estimate the real and imaginary parts of the complex ratio mask, λ r (k, l) and λ I (k, l) are the real and imaginary parts of cIRM.
[0075] Step S130: Obtain a weighted focusing covariance matrix.
[0076] Based on formula (4), the denoised and reverberated multi-channel enhanced signal can be obtained: In order to alleviate the degradation of DOA estimation performance caused by masking distortion of Z(k,l), a broadband spatial spectrum is used as an intermediate representation of the spatial orientation information of the sound source. The covariance matrix of the direct sound of the sound source can be estimated as follows:
[0077]
[0078] Wherein, j=0,1,...,J is the snapshot index, and J is the snapshot number.
[0079] Frequency focusing can avoid the high complexity and rank-deficient covariance matrix problems caused by multi-band calculations; at the same time, using the probability of speech existence to weight the sub-bandwidth focusing covariance matrix can reduce the impact of the noise-dominated frequency band; in addition, dividing the frequency band into sub-bands can reduce the focusing error in the entire bandwidth while improving the resolution of the spatial spectrum frequency dimension. Therefore, the present patent adopts a weighted frequency focusing method of sub-band segmentation to construct a broadband spatial spectrum. If the K frequency bands are evenly divided into W sub-bandwidths, frequency focusing can be performed within the W sub-bandwidths to obtain a weighted focusing covariance matrix, as shown in equations (7) to (8).
[0080]
[0081] in, is the weighted focusing covariance matrix of the w-th sub-bandwidth, w = 1, 2, ..., W is the sub-bandwidth index, f w is the focusing reference frequency of the wth sub-bandwidth, k w and K w are the band index and number of bands in the w-th sub-bandwidth, respectively, λ(k w ) is the estimated probability of direct speech, C(k w ) is the kth w The focusing matrix of the frequency bands.
[0082] Since the number of snapshots in practical applications is often limited (hopefully as small as possible to accommodate streaming algorithms and processing of moving sound sources), the statistical The sound source spatial characteristics covered are not robust, which will affect the subsequent estimated sub-bandwidth focused spatial spectrum characteristics, causing the spatial information of some weaker sound sources to be submerged by residual noise, residual reverberation and characteristics of other sound source signals. In order to improve this problem, the present disclosure adopts a sub-bandwidth overlapping method to obtain the sub-bandwidth focused spatial spectrum. The overlapping between sub-bandwidths can bring two advantages: first, the sub-bandwidth overlapping can increase the number of sub-bandwidths, thereby further improving the frequency dimension of the broadband focused spatial spectrum; second, the bandwidth of the sub-bandwidth can be increased without reducing the frequency resolution, so that richer frequency dimension features can be included. middle.
[0083] Step S140: Obtain broadband focused spatial spectrum features based on beamforming, and construct a convolutional neural network model.
[0084] Based on equation (7), some traditional methods can be used to estimate DOA information, such as MUSIC. However, since the number of sound sources is unknown, and masking and weighted focusing cannot directly eliminate the influence of residual noise and phase distortion, high-resolution subspace methods like MUSIC are very sensitive to the estimated noise subspace. Therefore, the present invention adopts the minimum variance distortionless response (MVDR) beamforming method to model the sub-bandwidth focused spatial spectrum, as shown in equations (9) to (10):
[0085]
[0086] in, is the spatial response power (Spatial Response Power, SRP) of the w-th sub-bandwidth, is the angle of spatial spectrum traversal, The reference frequency is f w At The direction of the steering vector, is the filter weight of the MVDR beamformer.
[0087] Based on the focused spatial spectra of W sub-bandwidths obtained by formula (9), effective broadband focused spatial spectrum features can be obtained without estimating the number of sound sources.
[0088] In order to improve the resolution and error tolerance of the algorithm DOA estimation, the data sets of equations (1) and (9) are combined, and the DOA spatial spectrum is mapped by focusing the spatial spectrum from the broadband frequency based on CNN. The convolutional neural network model structure is as follows: Figure 4 As shown in Figure 1, it mainly consists of two parts: the first part is a 4-layer convolutional layer for extracting identification features, and the second part is a 4-layer fully connected layer for mapping the DOA spatial spectrum. The input of CNN is a broadband focused spatial spectrum with a dimension of Υ×W. Each convolutional layer consists of 128 filters with a kernel of 2×1. The dimensions of the four fully connected layers are 512×1, 256×1, 128×1 and Υ×1 respectively. The output of the last fully connected layer is the estimated DOA spatial spectrum, where Υ is the number of DOA spatial spectrums. The mean square error (MSE) shown in formula (11) is used as the loss function for weight optimization of DOA estimation CNN:
[0089]
[0090] The DOA spatial pseudo spectrum of equation (1) is used as the real DOA spatial spectrum φ real ,Right now
[0091] Step S150: testing and reasoning the DOA spatial spectrum of the decoding process based on the mask estimation network model and the sound source DOA estimation network model.
[0092] Through steps S110 to S150, the trained DNN masking estimator and CNN spatial spectrum mapper can be obtained. The test reasoning process is as follows Figure 2 By finding the significant peak of the output DOA spatial spectrum, the corresponding DOA estimation result can be obtained.
[0093] In the embodiment of this example, the DOA estimation method based on broadband frequency focused spatial spectrum and convolutional neural network disclosed in the present invention is:
[0094] Joint masking is used to enhance the direct sound signal of each sound source in the multi-sound source scenario, and the wideband frequency focused spatial spectrum with multiple sub-bandwidths overlapped is used as the intermediate representation of the DOA estimation result, which reduces the impact of the nonlinear distortion caused by masking processing on the DOA estimation algorithm.
[0095] In the process of solving the broadband frequency focused spatial spectrum, a series of optimization treatments were carried out: first, the mask mean within the snapshot was used as the weight to weight the focused covariance matrix, thereby reducing the influence of the noise-dominated frequency band; second, MVDR beamforming was used to solve the sub-bandwidth spatial response power spectrum as the sub-bandwidth frequency focused spatial spectrum, thereby avoiding the problems of sound source number estimation and subspace decomposition error; third, the sub-bandwidth overlapping method was used to improve the frequency resolution and sub-bandwidth width of the broadband frequency focused spatial spectrum, thereby enhancing the robustness of the focused covariance matrix characteristics under finite snapshot.
[0096] The processing and training process of masking enhancement, broadband frequency focusing spatial spectrum estimation and CNN-based DOA spatial spectrum estimation were designed to improve the robustness and fault tolerance of the DOA estimation algorithm.
[0097] Compared with the prior art, the present disclosure proposes a method for estimating the direction of arrival of multiple sound sources based on wideband frequency-focused spatial spectrum and convolutional neural network by combining a signal modeling method with a neural network method. The method has the following beneficial effects: first, masking enhancement is combined with a wideband spatial spectrum based on frequency focusing, that is, a wideband frequency-focused spatial spectrum is used as an intermediate representation of the sound source spatial information, thereby alleviating the degradation of DOA estimation performance caused by the nonlinear distortion of masking enhancement; second, in order to improve the representation capability of the wideband frequency-focused spatial spectrum for the sound source spatial information, each sub-bandwidth is overlapped to improve the frequency resolution of the wideband spatial spectrum feature, and the mask mean is introduced to weight the focusing covariance matrix to improve the stability of the reverberation environment feature; third, by designing a DOA spatial spectrum label with a certain resolution, the DOA spatial spectrum feature is mapped by the wideband frequency-focused spatial spectrum feature through CNN, thereby improving the fault tolerance of the wideband spatial spectrum feature for the representation of the sound source spatial information, thereby obtaining robust DOA estimation performance in a complex acoustic environment.
[0098] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the present disclosure.
[0099] Reference Figure 5 , shows a structural block diagram of a multi-sound source arrival direction estimation device 200 based on frequency-focused spatial spectrum of the present disclosure, the device includes a data construction module 210, a model construction module 220, a covariance matrix acquisition module 230, a spatial spectrum training module 240 and a spatial spectrum reasoning module 250, wherein:
[0100] The data construction module 210 is used to generate a microphone array data set from a single sound source clean speech corpus using a preset simulation environment and a microphone array signal simulator, and to construct a training data set based on the microphone array data set.
[0101] The model building module 220 is used to generate a mask estimation network model by performing sound source masking training on the training data set; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking.
[0102] The covariance matrix acquisition module 230 is used to obtain the enhanced multi-sound source signal according to the mask estimation network model, and divide the enhanced multi-sound source signal into sub-bandwidths by frequency band overlapping to obtain the weighted focused covariance matrix of each sub-bandwidth.
[0103] The spatial spectrum training module 240 is used to construct a broadband frequency-focused spatial spectrum feature based on the sub-bandwidth weighted focusing covariance matrix obtained by beamforming, and to construct a convolutional neural network model based on the broadband frequency-focused spatial spectrum feature, and to estimate the multi-sound source DOA spatial spectrum using the convolutional neural network model.
[0104] The spatial spectrum inference module 250 is used to predict the sound source DOA spatial spectrum according to the decoding process of the mask estimation network model and the sound source DOA estimation network model, and complete the estimation of multiple sound source DOAs.
[0105] The multi-source arrival direction estimation device based on frequency-focused spatial spectrum processes multi-source signals and uses masking estimation, spatial spectrum training and reasoning to build an accurate multi-source localization and separation model. The device works together through multiple modules, from environmental data generation, network model training, to the final DOA estimation, gradually improving the accuracy and robustness of sound source localization and separation.
[0106] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0107] Optionally, an embodiment of the present disclosure further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0108] The embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0109] Figure 6 800 is a block diagram of an electronic device 800 shown in the present disclosure. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0110] Reference Figure 6, the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0111] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0112] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0113] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0114] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0115] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0116] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0117] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0118] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0119] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0120] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0121] Figure 7 19 is a block diagram of a computer-readable storage medium 1900 shown in the present disclosure. For example, the computer-readable storage medium 1900 may be provided as a server.
[0122] Reference Figure 7 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0123] The computer readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0124] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0125] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0126] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, but the present disclosure is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present disclosure, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present disclosure and the claims, all of which are within the protection of the present disclosure.
[0127] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present disclosure can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0128] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0129] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0130] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0131] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0132] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks or optical disks.
[0133] The above is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A method for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum, characterized in that: The method comprises: Using a preset simulation environment and a microphone array signal simulator, a microphone array data set is generated from a single sound source clean speech corpus, and a training data set is constructed based on the microphone array data set; By performing sound source masking training on the training data set, a mask estimation network model is generated; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking; According to the mask estimation network model, an enhanced multi-sound source signal is obtained, and the enhanced multi-sound source signal is divided into sub-bandwidths by frequency band overlapping to obtain a weighted focusing covariance matrix of each sub-bandwidth; Based on the weighted focusing covariance matrix of each sub-bandwidth, a sub-bandwidth frequency focusing covariance matrix obtained by beamforming is used to construct a broadband frequency focusing spatial spectrum feature, and a convolutional neural network model is constructed based on the broadband frequency focusing spatial spectrum feature, and the convolutional neural network model is used to estimate the DOA spatial spectrum of multiple sound sources; The mask estimation network model and the multi-sound source DOA spatial spectrum estimation network model are used to decode and obtain the multi-sound source DOA spatial spectrum, and significant peaks are extracted to complete the estimation of the multi-sound source DOA.
2. The method according to claim 1, characterized in that The microphone array data set includes: clean microphone array speech signals, noisy and reverberant microphone array signals, and sound source DOA labels.
3. The method according to claim 2, characterized in that Constructing a training data set based on the microphone array data set, including: Using the simulation environment, constructing masking training data by performing short-time Fourier transform on the clean microphone array speech signal and the noisy and reverberant microphone array signal; According to the short-time Fourier transform structure of the noisy and reverberant microphone array signal, the activity area of the clean microphone array speech is detected and labeled by an endpoint check algorithm to construct sound source DOA spatial spectrum training data; A training data set is generated by integrating the masking training data and the sound source DOA spatial spectrum training data.
4. The method according to claim 1, characterized in that The enhanced multi-sound source signals are divided into sub-bandwidths by means of frequency band overlap, including: The broadband spatial spectrum is obtained by focusing the covariance matrix of the sub-bandwidth as an intermediate representation of the spatial orientation information of the sound source; The signal bandwidth is divided into multiple sub-bandwidths in a sub-bandwidth overlapping manner, and the speech existence probability of each sub-bandwidth is obtained based on a speech existence probability algorithm, and the signals of each sub-bandwidth are weighted to obtain a sub-bandwidth weighted focusing covariance matrix.
5. The method according to claim 1, characterized in that The sub-bandwidth frequency focusing covariance matrix obtained by beamforming constructs broadband frequency focusing spatial spectrum features, including: The spatial spectrum of the sub-bandwidth is modeled based on the minimum variance distortion-free response beamforming method, and the sub-bandwidth frequency-focused spatial spectrum characteristics based on beamforming are obtained.
6. The method according to claim 1, characterized in that The convolutional neural network comprises: The convolutional neural network includes a convolutional layer for extracting identification features and a fully connected layer for mapping the DOA spatial spectrum of the sound source wave.
7. The method according to claim 1, characterized in that Using the mask estimation network model and the multi-sound source DOA spatial spectrum estimation network model, decoding to obtain the multi-sound source DOA spatial spectrum includes: Based on the decoding process of the mask estimation network model and the sound source DOA estimation network model, the sound source wave DOA prediction result is generated by calculating the significant peaks of the multi-sound source DOA spatial spectrum; Based on the sound source wave DOA prediction result, the DOA of multiple sound sources is estimated.
8. A device for estimating the direction of arrival of multiple sound sources based on frequency-focused spatial spectrum, characterized in that: The device comprises: A data construction module, used to generate a microphone array data set from a single sound source clean speech corpus using a preset simulation environment and a microphone array signal simulator, and to construct a training data set based on the microphone array data set; A model building module, used to generate a mask estimation network model by performing sound source masking training on the training data set; the loss function of the mask estimation network model is the mean square error of the sound source complex ideal ratio masking; A covariance matrix acquisition module, used to acquire the enhanced multi-sound source signal according to the mask estimation network model, and divide the enhanced multi-sound source signal into sub-bandwidths by frequency band overlap to acquire a weighted focusing covariance matrix of each sub-bandwidth; A spatial spectrum training module, for constructing broadband frequency-focused spatial spectrum features based on the weighted focusing covariance matrix of each sub-bandwidth and the sub-bandwidth frequency-focused covariance matrix obtained by beamforming, and constructing a convolutional neural network model based on the broadband frequency-focused spatial spectrum features, and estimating the DOA spatial spectrum of multiple sound sources using the convolutional neural network model; The spatial spectrum inference module is used to predict the sound source DOA spatial spectrum according to the decoding process of the mask estimation network model and the sound source DOA estimation network model, and complete the estimation of multiple sound source DOAs.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-sound-source arrival direction estimation method and device and nonvolatile storage medium
CN120630101A