Adaptive microphone array separation enhancement method and system
By employing the cGMM adaptive microphone array algorithm and the deep Bayesian source separation method, the problems of speaker position changes and frequency arrangement ambiguity in microphone array technology are solved, achieving efficient speech enhancement and separation in complex scenarios.
Patent Information
- Application Number
- CN202211712017.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing microphone array technology has difficulty maintaining a fixed speaker orientation in complex scenarios, resulting in poor speech enhancement effects. Furthermore, single-channel separation is prone to problems with speech separation due to ambiguity in direction.
The cGMM adaptive microphone array algorithm is adopted, which collects signals through an omnidirectional microphone array, estimates the time-frequency domain mask and direction of arrival using the spatial covariance matrix, and performs joint training in conjunction with the deep Bayesian source separation method to solve the frequency arrangement ambiguity and achieve speech enhancement and separation.
It effectively enhances speech in complex scenarios, resolves the impact of multiple speaker position changes, improves speech recognition accuracy, effectively removes noise, and enhances speech quality in complex environments.
Smart Images

Figure CN115802245B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of array microphone technology, and more specifically to an adaptive microphone array separation enhancement method and system. Background Art
[0002] Currently, microphone array technology is widely used in voice conversation scenarios. Array algorithms such as angle localization, voice enhancement, and voice separation are typically used to enhance the voice in the target dialogue direction in order to improve the recording signal-to-noise ratio and the accuracy of voice recognition.
[0003] Speech signal enhancement methods picked up by microphone signals typically use array angle localization to determine the target direction angle, and then use beamforming methods to improve the signal-to-noise ratio at that angle. However, in real-world scenarios, it is usually difficult to keep the speaker in a fixed direction, and the speaker may change position within a single dialogue. Furthermore, in single-channel separation applications, complex scenarios often result in the separation of speech with ambiguous directions. Summary of the Invention
[0004] This invention provides an adaptive microphone array separation enhancement method, which adopts the cGMM adaptive microphone array algorithm. It can not only complete the enhancement task in software, but also design intelligent electronic products to realize the function of enhancing target dialogue in various dialogue scenarios.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an adaptive microphone array separation enhancement method, comprising the following scheme:
[0006] Step 1: Multi-channel signal mixing; Audio is acquired through an array of omnidirectional microphones and an acquisition module composed of electronic components, and converted into m-channel digital signals with time t through an analog-to-digital converter;
[0007] Step 2: Separate and locate the network; estimate the time-frequency domain TF mask and DoAs of each source by mixing spatial covariance matrices (SCMs);
[0008] Step 3: Deep Bayesian source separation; the separation and localization network is trained using multi-channel mixed signals, and the frequency arrangement ambiguity problem is solved within a unified framework.
[0009] Preferably, step one further includes performing a Fast Fourier Transform (FFT) on the two-dimensional digital signal to transform it into a time-frequency domain signal, where the frequency domain dimension is defined as f.
[0010] Preferably, in step two, the objective function is derived as the lower evidence limit ELBO of the cGMM with a TF mask and DoAs as latent variables.
[0011] Preferably, in step three, the training is based on a latent LDA Dirichlet distribution model, which uses the source TF mask and DoAs as latent variables, and the objective function is the ELBO of the spatial model, which consists of the expectation of the likelihood function and the KL divergence between the network output and its prior distribution.
[0012] Preferably, the TF mask and DoAs of the potential sound source are jointly estimated, and the observable multi-channel spectrogram x is obtained. tf Represented as a spectrum diagram of K sound sources s tfk The sum, that is:
[0013]
[0014] in:
[0015] t is the time frame;
[0016] f is the frequency bin;
[0017] m is the microphone index;
[0018] k is the source index;
[0019] d is the direction index;
[0020] z tfk ∈{0,1} is a TF mask indicating which sound source is correlated in each TF bin, w kd ∈{0,1} is a DoA variable, which assigns the sound source k to a DoA candidate d∈{1,…,D}, a fd It is a direction-guided vector.
[0021] Preferably, ELBO maximization corresponds to minimizing the KL divergence between the variational distribution and the true posterior distribution; the ELBO update method is as follows: first predict the TF mask and DoAs for each mixed recording, then update the model parameters, and finally use the stochastic gradient descent (SGD) method to calculate and update the network parameters.
[0022] Preferably, the trained network initializes the TF mask through network output to improve the performance of the multi-channel EM algorithm, which is used to solve the expected maximization algorithm EM of cGMM in the subsequent e-step and m-step alternating iterations; the TF mask z is updated in the e-step. tfk and DoAsw kd Maximize ELBO L; update parameters in m steps.
[0023] Preferably, a separation network g is used. tfk Output to TF mask z tfk Perform initialization, that is:
[0024] Preferably, an adaptive microphone array separation enhancement system is used for an adaptive microphone array separation enhancement method, characterized in that it includes: a multi-channel conversation signal acquisition unit, a multi-channel unsupervised training module, and speech enhancement; the multi-channel conversation signal acquisition unit, the separated speech obtained by multi-channel unsupervised training, and the single-channel enhancement can completely remove noise and become normal single-person speech through speech enhancement.
[0025] The beneficial effects of this invention are as follows: It does not rely on angle positioning; the invention uses only location information for speech enhancement, and changes in the dialogue position do not affect the enhancement effect, thus solving the problem of positional variations in multi-speaker scenarios. It also effectively utilizes multi-channel correlation, using multi-channel adaptive hybrid model information for more effective and accurate separation, resolving scenarios where multiple speakers are speaking simultaneously and enriching the dialogue content. It resolves the arrangement ambiguity that occurs after single-channel separation, is unaffected by noise and reverberation environments, and can effectively enhance speech in complex scenarios. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the multi-channel unsupervised training process of the present invention;
[0028] Figure 2 This is a schematic diagram of the enhanced system architecture of the present invention. Detailed Implementation
[0029] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Traditional neural decoupling methods require a large amount of supervised data to achieve good performance. Although spatially based multichannel methods can work without such training data, they are often very sensitive to parameter initialization and degrade when the sources are close to each other.
[0031] according to Figure 1 , Figure 2 As shown, the present invention provides an adaptive microphone array separation enhancement method, including the following scheme:
[0032] Step 1: Multichannel mixture signal:
[0033] The mixed signal in this invention is a multi-channel microphone speech time-frequency domain signal. It is first obtained using an array of various omnidirectional microphones with good consistency and an acquisition module composed of electronic components. The signal is then converted into m-channel digital signals with time t via an analog-to-digital converter. Next, the two-dimensional digital signal is transformed into a time-frequency domain signal using FFT (Fast Fourier Transform) (in this invention, the frequency domain dimension is defined as f).
[0034] Step 2: Separation network and localization network:
[0035] This invention addresses frequency ordering ambiguity by jointly training an unsupervised neural source separation and localization network, rather than using the correlation of a mask. A common method for separating multi-channel mixed signals is to mask each time-frequency point. This mask is typically estimated through manual feature clustering at each time-frequency point. Because these models are built independently at each frequency point, they suffer from frequency ordering ambiguity. To address this issue, this invention estimates the time-frequency mask (TF mask) and DoAs for each source by using pre-measured direction-of-arrival (DoAs) vectors to characterize potential directions of arrival (DoAs).
[0036] The objective function of this invention is derived as the Evidence Lower Bound (ELBO) of a cGMM with a TF mask and DoAs as latent variables. Given the geometry of a microphone array, this invention trains two networks to estimate the post-probabilities of the TF mask and DoAs, respectively. Since DoAs can be used to compute the number of sources in a mixed recording, our framework can be extended to process training data containing an unknown number of sources using a parameter-free Bayesian model.
[0037] Step 3: Deep Bayesian Source Separation (DBSS):
[0038] This invention utilizes only multi-channel mixed signals to train a separation and localization network, and solves the frequency arrangement ambiguity problem within a unified framework. The training is based on a latent Dirichlet allocation (LDA) model, which uses the source's TF mask and DoAs as latent variables. The objective function is an ELBO of the spatial model, consisting of the expectation of the likelihood function and the Kullback-Leibler (KL) divergence between the network output and its prior distribution.
[0039] To jointly estimate the TF mask and DoAs of potential sound sources, an observable multichannel spectrogram x tf Represented as a spectrum diagram of K sound sources s tfk The sum of:
[0040]
[0041] Where ztfk∈{0,1} is a TF mask, indicating which sound source is correlated in each TF bin, w kd ∈{0,1} is a DoA variable, assigning the sound source k to a DoA candidate d∈{1,…,D}, where afd is the direction steering vector. In the invention, if we assume the potential direction d is a direction with a 5° intersecting angle on the horizontal plane, then d=360 / 5=72.
[0042] ELBO maximization corresponds to minimizing the KL divergence between the variational distribution and the true posterior distribution. This framework iteratively updates the parameters alternately until convergence. Since analyzing and computationally analyzing these variables is also difficult, we use ELBO to update them, specifically as follows:
[0043] 1) Predict the TF mask and DoAs for each mixed recording;
[0044] 2) The model parameters will be updated;
[0045] 3) Calculate and update network parameters using the stochastic gradient descent (SGD) method.
[0046] While a trained network can be used to separate resources from mono-mixed signals, it can also improve the performance of the multi-channel EM algorithm by initializing the TF mask using the network output. The Expectation Maximization (EM) algorithm used to solve the cGMM iterates alternately between subsequent e-steps and m-steps. The TF mask z is updated at the e-step. tfk and DoAs w kd This maximizes ELBO L; on the other hand, m steps use updated parameters.
[0047] Because the EM algorithm updates these variables alternately before convergence, this invention employs cautious initialization to avoid getting trapped in local optima. It utilizes a split network g. tfk Output to TF mask z tfk Initialization is performed. This is because the positioning network g... tfk Spatial bias that may lead to overfitting of training data is addressed by using the following formula instead of the output of the localization network hkd to initialize DoA w in this invention. kd :
[0048]
[0049] In summary, this method uses a cost function based on a spatial model of complex Gaussian Mixture Model (cGMM). This model uses the time-frequency mask of the source and the direction of arrival as latent variables to train the separation and localization networks, and estimates these variables respectively.
[0050] This joint training addresses the frequency arrangement ambiguity of spatial models within a unified deep Bayesian framework. Furthermore, the pre-trained network can be used not only for single-channel separation but also effectively initialize multi-channel separation algorithms. Experimental results on simulated speech mixing demonstrate that this method outperforms traditional initialization methods.
[0051] Therefore, this method does not rely on angle positioning. This invention uses only location information for speech enhancement, and changes in dialogue position do not affect the enhancement effect, thus solving the problem of positional variations in multi-speaker scenarios. It also effectively utilizes multi-channel correlation, employing a multi-channel adaptive hybrid model for more effective and accurate separation, addressing scenarios with multiple speakers simultaneously and enriching the dialogue content. It resolves the arrangement ambiguity that occurs after single-channel separation, is unaffected by noise and reverberation, and can effectively enhance speech in complex scenarios.
[0052] As an implementation scheme, a multi-channel conversation enhancement system includes a multi-channel conversation signal acquisition unit, multi-channel unsupervised training-derived separated speech, and single-channel enhancement. The separated speech signal may still have residual spatial reverberation and noise pollution sources, which can be completely removed by speech enhancement to become normal single-person speech.
[0053] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An adaptive microphone array separation enhancement method, characterized in that... The following solutions are included: Step 1: Multi-channel signal mixing; Audio is acquired through an array of omnidirectional microphones and an acquisition module composed of electronic components, and converted into m-channel digital signals with time t through an analog-to-digital converter; Step 2: Separate and locate the network; estimate the time-frequency domain TF mask and DoAs of each source by mixing spatial covariance matrices (SCMs); Step 3: Deep Bayesian source separation; a separation and localization network is trained using multi-channel mixed signals, and the frequency arrangement ambiguity problem is solved within a unified framework; Step one also includes performing a Fast Fourier Transform (FFT) on the two-dimensional digital signal to transform it into a time-frequency domain signal, where the frequency domain dimension is defined as f. In step two, the objective function is derived as the lower evidence limit ELBO of cGMM with TF mask and DoAs as latent variables. In step three, the training is based on a latent LDA Dirichlet distribution model, which uses the source TF mask and DoAs as latent variables. The objective function is the ELBO of the spatial model, which consists of the expectation of the likelihood function and the KL divergence between the network output and its prior distribution.
2. The adaptive microphone array separation enhancement method according to claim 1, characterized in that: Joint estimation of TF mask and DoAs of potential sound sources, observable multi-channel spectrogram x tf Represented as a spectrum diagram of K sound sources s tfk The sum, that is: in: t is the time frame; f is the frequency bin; k is the source index; d is the direction index; z tfk ∈{0,1} is a TF mask, indicating which sound source is correlated in each TF bin, w kd ∈{0,1} is a DoA variable, which assigns the sound source k to a DoA candidate d∈{1,…,D}, a fd It is a direction-guiding vector.
3. The adaptive microphone array separation enhancement method according to claim 2, characterized in that: ELBO maximization corresponds to minimizing the KL divergence between the variational distribution and the true posterior distribution. The ELBO update method is as follows: first predict the TF mask and DoAs for each mixed recording, then update the model parameters, and finally use the stochastic gradient descent (SGD) method to calculate and update the network parameters.
4. The adaptive microphone array separation enhancement method according to claim 3, characterized in that: The trained network is initialized with a TF mask through network output to improve the performance of the multi-channel EM algorithm, which is used to solve the expected maximization algorithm EM alternating iterations in the subsequent e-steps and m-steps of cGMM. eStep updates TF mask z tfk and DoAs w kd Maximize ELBO L; update parameters in m steps.
5. The adaptive microphone array separation enhancement method according to claim 4, characterized in that: Using the separation network g tfk Output to TF mask z tfk Perform initialization, that is:
6. An adaptive microphone array separation enhancement system, used in the adaptive microphone array separation enhancement method of claim 1, characterized in that, include: Multi-channel conversation signal acquisition unit, multi-channel unsupervised training module, and speech enhancement; The multi-channel conversation signal acquisition device, the separated speech obtained from multi-channel unsupervised training, and the single-channel enhancement completely remove noise and turn it into normal single-person speech through speech enhancement.
Citation Information
Patent Citations
Speech enhancement method and device based on dual-channel neural network time-frequency masking, and hearing-aid equipment
CN114078481A
System and apparatus for tracking moving audio sources
WO2017129239A1