Speech enhancement and separation method and device based on directional convolution beam forming
By adopting a directive convolution beamforming method in speech enhancement technology, combining Kalman filtering and multi-channel linear prediction, the problems of early reverb removal and time-varying variance sensitivity in the prior art are solved, and efficient speech signal enhancement and separation effects are achieved.
Patent Information
- Application Number
- CN202411981375.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The prior art is difficult to effectively remove early reverb and sensitivity to time-varying variance in speech enhancement, resulting in distortion of target sound source and residual interference sound source.
The Kalman gain linear prediction error model is constructed through Kalman filtering and multi-channel linear prediction, and the minimum variance-free distortion-response beamforming model is constructed using the maximum directional beamformer and the maximum zero notch beamformer. The time-varying variance is combined to achieve enhanced and separated speech signals.
Effectively remove early reverb, improve the clarity and nature of speech signals, and achieve robust denoising and speech separation performance by estimating the noise covariance matrix in real time.
Smart Images

Figure CN119943085A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of acoustic signal processing, and in particular to a speech enhancement and separation method and device based on directional convolution beamforming. Background Art
[0002] In speech enhancement based on microphone arrays, beamforming technology is one of the commonly used filtering methods. Its main goal is to improve signal quality according to the spatial distribution of acoustic components such as sound source, noise and reverberation through spatial filtering technology, and to achieve denoising, dereverberation and sound source separation. Classic beamformers, such as delayed sum beamformer, minimum power distortionless response (MPDR) beamformer, minimum variance distortionless response (MVDR) beamformer, etc., perform well in noise suppression, but are weak in dereverberation. This problem has triggered a large number of studies based on multi-channel linear prediction, especially the weighted power error (WPE) algorithm. The WPE algorithm achieves effective dereverberation performance by modeling the linear prediction relationship of multi-channel time series signals; therefore, many studies in the past cascaded the WPE algorithm with the beamformer to achieve better denoising and dereverberation performance. However, WPE and beamformer usually work independently of each other, which makes it impossible to achieve the optimal effect when jointly optimized.
[0003] In the prior art, multi-channel linear prediction methods based on convolutional beamformers (CBFs) are proposed, such as weighted power minimum distortionless response (WPD) and convolutional beamformers based on source exact decomposition (SWF-CBF). These methods combine beamforming and multi-channel linear prediction for dereverberation, and aim to obtain better denoising and dereverberation effects by jointly optimizing the cost function. Although the technology based on convolutional beamformers has made certain progress, there are still some problems that cannot be ignored. First, multi-channel linear prediction methods are usually only applicable to the removal of late reverberation, which leads to the fact that in practical applications, early reverberation components often remain in the separated sound source signals, affecting the clarity and naturalness of the speech. Secondly, many convolutional beamformers are implemented based on MPDR beamformers, and the performance of convolutional beamformers implemented based on MPDR often depends on the accuracy of time-varying variance. In practical environments, the estimation of time-varying variance often has errors, which will lead to distortion of the target sound source and residual interference sound sources.
[0004] Therefore, although the joint optimization of the convolutional beamformer can improve the dereverberation performance to a certain extent, it still faces great challenges in reducing early reverberation and sensitivity to time-varying variance, and an effective solution is urgently needed. Summary of the invention
[0005] The present disclosure shows a method and device for speech enhancement and separation based on directional convolution beamforming.
[0006] In a first aspect, the present disclosure provides a method for speech enhancement and separation based on directional convolution beamforming, the method comprising:
[0007] According to the microphone array with a preset number of array elements, a multi-channel observation signal is obtained, and a historical observation signal matrix suitable for the multi-channel linear prediction dereverberation task is constructed;
[0008] According to the historical observation signal matrix, a Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction;
[0009] The maximum directivity beamformer and the maximum nulling beamformer are used to construct a minimum variance distortion-free response beamforming model based on directivity gain.
[0010] Estimate speech and noise covariance matrices based on a directional minimum variance distortionless response beamformer, and estimate time-varying variance based on separated speech sources;
[0011] The time-varying variance is used to jointly establish the Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model to establish a directional convolution beamforming model, and the enhancement and separation of multi-source speech signals containing noise and reverberation are completed through alternating iterations.
[0012] In an exemplary embodiment of the present disclosure, a multi-channel observation signal is obtained, and a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed, including:
[0013] Acquire multi-channel observation signals of a preset number of speech sources according to a microphone array with a preset number of array elements;
[0014] Performing short-time Fourier transform processing on the multi-channel observation signal;
[0015] Based on the multi-channel observation signals processed by short-time Fourier transform, a historical observation signal matrix suitable for the multi-channel linear prediction dereverberation task is constructed.
[0016] In an exemplary embodiment of the present disclosure, a Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction, including:
[0017] According to the historical observation signal matrix, based on Kalman filtering, a multi-channel Kalman filter model is established;
[0018] Based on the multi-channel Kalman filter model, a multi-channel linear prediction model is jointly established to establish a dereverberation construction.
[0019] In an exemplary embodiment of the present disclosure, a maximum directivity beamformer and a maximum nulling beamformer are used to construct a minimum variance distortion-free response beamforming model based on directivity gain, including:
[0020] According to the historical observation signal matrix, a dereverberation signal is estimated using a multi-channel linear prediction model based on Kalman filtering;
[0021] estimating a denoised and separated speech signal using a directional beamformer based on the dereverberated signal;
[0022] Based on the Kalman gain linear prediction error model, a noise covariance matrix and an expected signal covariance matrix are established by splitting the speech plus noise covariance matrix;
[0023] Based on the dereverberation signal, a minimum variance distortionless response beamforming model is constructed by estimating a target signal covariance matrix and a non-target signal covariance matrix based on a maximum directivity beamformer and a maximum null beamformer;
[0024] Based on the constructed minimum variance distortion-free response beamformer, the separated signals of each speech source are obtained, and the covariance matrix of the speech signal is constructed;
[0025] Based on the dereverberation signal covariance matrix, the noise covariance matrix is obtained by subtracting the speech signal covariance matrix.
[0026] In an exemplary embodiment of the present disclosure, the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model are jointly constructed, including:
[0027] According to the dereverberation signal, beamforming weights are calculated by using a time-varying variance-weighted covariance matrix to establish a convolutional beamforming model;
[0028] Based on the Lagrange multiplier method, the minimum variance beamforming model is solved, and an equivalent transformation is performed on the solution of the beamforming model;
[0029] The covariance matrix of non-target signals and the covariance matrix of target signals in the solution of the equivalently transformed beamforming model are estimated.
[0030] In an exemplary embodiment of the present disclosure, estimating a non-target signal covariance matrix and a target signal covariance matrix in a solution of a beamforming model subjected to an equivalent transformation includes:
[0031] The speech source noise covariance matrix is estimated according to the silent segment, and the correlation items between signal components are eliminated based on the steering vector joint matrix to establish a maximum null notch beamformer;
[0032] According to the directional steering vector, the integration is performed in the directional space and combined with the diagonal loading matrix to establish a maximum directivity beamformer;
[0033] Based on the maximum directivity beamformer and the maximum nulling beamformer, estimating the masking vector by directivity gain;
[0034] Establishing a general model for estimating non-target signals by estimating the masking vector;
[0035] Based on the recursive form, the estimation of the non-target signal covariance matrix and the target signal covariance matrix is completed through the estimation general model.
[0036] In an exemplary embodiment of the present disclosure, the enhancement and separation of the noisy and reverberant multi-source speech signals are completed by an alternating iterative method, including:
[0037] Initialize parameters and estimate the steering vector;
[0038] According to the estimated steering vector, a dereverberation filter is estimated to obtain a dereverberation signal;
[0039] Obtaining uniform beamforming weights by estimating directivity gain beamforming matrix weights;
[0040] Iteratively updating the speech plus noise covariance matrix and the time-varying variance based on the unified beamforming weights;
[0041] By alternately iterating the multi-channel linear prediction based on Kalman filtering and the beamformer based on directivity gain, the performance of convolution beamforming is achieved, and the enhancement and separation of speech signals are completed.
[0042] In a second aspect, the present disclosure discloses a speech enhancement and separation device based on directional convolution beamforming, the device comprising:
[0043] An observation signal construction module is used to obtain multi-channel observation signals according to a microphone array with a preset number of array elements, and to construct a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task;
[0044] An error modeling module based on Kalman filtering is used to construct a Kalman gain linear prediction error model based on the historical observation signal matrix using Kalman filtering and multi-channel linear prediction;
[0045] Directivity beamforming modeling module, used to combine the maximum directivity beamformer and the maximum nulling beamformer to build a minimum variance distortion-free response beamforming model based on directivity gain;
[0046] A module for estimating a speech plus noise covariance matrix and a time-varying variance, for estimating a speech and noise covariance matrix based on a minimum variance distortionless response beamformer of a directivity gain, and estimating a time-varying variance based on a separated speech source;
[0047] The alternating iterative processing module is used to utilize the time-varying variance to jointly establish the Kalman gain-based linear prediction error model and the minimum variance distortion-free response beamforming model based on the directivity gain, establish a directional convolution beamforming model, and enhance and separate the noisy and reverberant multi-source speech signals through alternating iteration.
[0048] In a third aspect, the present disclosure shows an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.
[0049] In a fourth aspect, the present disclosure shows a non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as described in any of the above aspects.
[0050] In a fifth aspect, the present disclosure shows a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any of the above aspects.
[0051] The present invention discloses a speech enhancement and separation method based on directional convolution beamforming. First, according to a microphone array with a preset number of array elements, a historical observation signal matrix is constructed by acquiring multi-channel observation signals; secondly, according to the historical observation signal matrix, a Kalman gain linear prediction error model based on Kalman filtering and multi-channel linear prediction dereverberation is constructed; then, based on a maximum directivity beamformer and a maximum nulling beamformer, a minimum variance distortionless response beamforming model based on directional gain is constructed; then, based on a minimum variance distortionless response beamformer based on directional gain, a speech and noise covariance matrix is estimated, and a time-varying variance is estimated based on a separated speech source; finally, by using the time-varying variance, a multi-channel linear prediction error model based on Kalman filtering and a minimum variance distortionless response beamforming model based on directional gain are jointly established to establish a directional convolution beamforming model, and the enhancement and separation of speech signals are completed through an alternating iteration.
[0052] The present invention utilizes Kalman filtering in conjunction with multi-channel linear prediction to achieve dereverberation, which can better achieve dereverberation in a noisy and reverberant environment. At the same time, a maximum directivity beamformer and a maximum nulling beamformer are used to construct a directivity gain, and the noise covariance matrix is estimated in real time to achieve robust denoising and speech separation performance. In addition, further, a multi-channel linear prediction model based on Kalman filtering and a minimum variance distortion-free response beamforming model based on directivity gain are combined through time-varying variance weighting to construct a directional gain convolution beamformer, which has robust denoising, dereverberation and speech separation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flowchart of the steps of a speech enhancement and separation method based on directional convolution beamforming disclosed in the present invention.
[0054] Figure 2 It is an algorithm principle diagram of a speech enhancement and separation method based on directional convolution beamforming disclosed in the present invention.
[0055] Figure 3 It is a structural block diagram of a speech enhancement and separation device based on directional convolution beamforming disclosed in the present invention.
[0056] Figure 4 is a block diagram of an electronic device disclosed herein.
[0057] Figure 5 is a block diagram of a computer-readable storage medium of the present disclosure. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0059]
Term explanation
[0060] Convolutional Beamformer (CBF) refers to a spatial filter constructed by multi-channel linear prediction and beamforming based on multi-channel observation signals at past moments. It is often used in denoising, dereverberation and interference sound source suppression of speech signals.
[0061] The minimum variance distortionless response (MVDR) beamformer is an adaptive beamformer that is constructed by minimizing the output power of non-target signals while constraining the target direction to be distortion-free. It is often used in the fields of enhanced processing of sonar, radar, communication, voice and other signals.
[0062] Multi-channel Linear Prediction (MCLP) refers to a technique that uses multi-channel observation signals at past moments in the form of linear prediction filtering to estimate the true value signal at the current moment. It is often used in late reverberation suppression of speech signals.
[0063] Reference Figure 1 , shows a flowchart of the steps of a speech enhancement and separation method based on directional convolution beamforming disclosed in the present invention, the method can be applied to electronic devices, wherein the method specifically may include the following steps:
[0064] Step S110, obtaining a multi-channel observation signal according to a microphone array with a preset number of array elements, and constructing a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task;
[0065] Step S120, constructing a Kalman gain linear prediction error model based on the historical observation signal matrix using Kalman filtering and multi-channel linear prediction;
[0066] Step S130, constructing a minimum variance distortion-free response beamforming model based on directivity gain by using a maximum directivity beamformer and a maximum nulling beamformer;
[0067] Step S140, estimating speech and noise covariance matrices based on a directional minimum variance distortionless response beamformer, and estimating time-varying variance based on the separated speech sources;
[0068] Step S150, using the time-varying variance, the Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model based on the directivity gain are jointly established to establish a directional convolution beamforming model, and the enhancement and separation of the noisy and reverberant multi-source speech signals are completed through an alternating iteration.
[0069] The speech enhancement and separation method based on directional convolution beamforming disclosed in the present invention combines Kalman filtering with multi-channel linear prediction, which can more effectively remove reverberation in a noisy and reverberant environment and improve the dereverberation performance. Secondly, the directional gain model constructed based on the maximum directional beamformer and the maximum null beamformer can estimate the noise covariance matrix in real time to achieve robust denoising and speech separation. In addition, the multi-channel linear prediction based on Kalman filtering is combined with the minimum variance distortion-free response beamforming technology based on directional gain through time-varying variance weighting to construct a directional gain convolution beamformer, which improves the overall robustness of denoising, dereverberation and speech separation. These advantages make this method have stronger processing capabilities in complex noise and reverberation environments.
[0070] In this exemplary embodiment, if Figure 2 As shown, the joint denoising, dereverberation and interference speech suppression method based on the directional convolution beamformer disclosed in the present invention is implemented by the following technical solutions:
[0071] In step S110, a multi-channel observation signal is acquired according to a microphone array with a preset number of array elements, and a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed.
[0072] First, assuming that the number of microphone array elements is M and the number of speech sources is Q, the multi-channel observation signal x(t) of the preset speech source can be obtained, as shown in formula (1):
[0073]
[0074] Where q=1,2,...,Q is the speech source index, is the observation signal collected by the microphone array at the historical time t, is the direct sound signal of the qth speech source, is the reflected sound signal (reverberation) of the qth speech source, is the noise collected by the microphone array at the historical time t, Represents real-time space.
[0075] Then, short-time Fourier transform (STFT) is performed on equation (1) to obtain equation (2):
[0076]
[0077] in, and are x(t) and s respectively. q (t), r q The STFT results of n(t) and n(t), k and l are the indices of frequency and frame number respectively, Represents complex space.
[0078] Based on formula (2) after STFT processing, the historical observation signal matrix suitable for multi-channel linear prediction dereverberation task is constructed:
[0079]
[0080] in, is the stacked signal of historical observation signals, b is the linear prediction delay, L g is the linear prediction filter order, the superscript " T ” indicates a transpose operation.
[0081] In step S120, a Kalman gain linear prediction error model is constructed based on the historical observation signal matrix using Kalman filtering and multi-channel linear prediction.
[0082] First, by setting the multi-channel linear prediction filter as Where L G =ML g , then the Kalman gain linear prediction error model can be modeled as follows:
[0083]
[0084] in, is the estimated Kalman linear prediction dereverberation model, E{} denotes the expected operation, and the superscript “ H ” indicates the conjugate transpose operation.
[0085] However, since G(k,l) is unknown in practice, Kalman filtering is used to Perform estimation calculations, such as equations (5)-(11):
[0086]
[0087] in, for A priori estimate of and I M The dimensions are L G ×L G and the M×M unit matrix, is the difference between the observed signal and the linear prediction reverberation component, for The sparse matrix of represents the Kronecker product, is the multi-channel Kalman gain, is the process noise covariance matrix, is the estimated speech plus noise covariance matrix.
[0088] In step S130, a minimum variance distortionless response beamforming model based on directivity gain is constructed by using a maximum directivity beamformer and a maximum null distortionless response beamformer.
[0089] In step S140, the speech and noise covariance matrices are estimated based on the directivity minimum variance distortionless response beamformer, and the time-varying variance is estimated based on the separated speech sources.
[0090] The Kalman gain K(k, l) shown in Equation (9) is a key parameter in Kalman filtering. Among them, in single-channel Kalman filtering, the Kalman gain can be regarded as: error signal × filter error / (error signal power + desired signal power), which is used to correct the filtering error at the current frequency point; while in multi-channel Kalman filtering, the Kalman gain becomes: error signal vector × filter error / (error signal covariance matrix + desired signal covariance matrix), which is used to correct the filtering error of the current frequency point vector. In single-channel processing, in order to improve the accuracy of the Kalman gain, many studies focus on the estimation of the desired signal power. In the present disclosure, the speech plus noise covariance matrix is split into a noise covariance matrix and a desired signal covariance matrix in two parts, and certain methods are respectively used for estimation.
[0091] First step, let the dereverberated signal be Z(k, l) ∈ C M×1 , and it can be estimated based on as follows:
[0092]
[0093] Second step, let the number of speech sources be Q (Q < M), then the beamforming weight of the q-th speech source is Then the denoised, dereverberated and separated signal can be estimated as follows:
[0094]
[0095] Third step, based on Equations (12) and (13), the noise covariance matrix Φ n (k, l) and the desired signal covariance matrix Φ s (k, l) can be estimated as follows:
[0096]
[0097] Among them, is the estimated power of the q-th speech source.
[0098] In the fourth step, since it takes Q beamformers to process each speech source one by one according to equation (13), in order to make the beamformer separate each speech source uniformly, the joint beamforming weight matrix is It is expressed as follows:
[0099]
[0100] Among them, the unified separated signal It can be expressed as:
[0101]
[0102] Among them, disg{} means taking the main diagonal elements.
[0103] The fifth step is to substitute equation (18) into equation (14) and equation (15) to obtain the noise covariance matrix Φ n (k,l), the estimation result of the expected signal covariance matrix Φs(k,l).
[0104] In step S150, the Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model are combined using time-varying variance to establish a directional convolution beamforming model, and the enhancement and separation of the noisy and reverberant multi-source speech signals are completed through alternating iterations.
[0105] First, based on the dereverberation signal Z(k,l), the beamforming weights are calculated through the time-varying variance-weighted covariance matrix to establish a beamforming model:
[0106]
[0107] in, and are the time-varying variance weighted covariance matrix and steering vector of the non-target signal of the qth speech source, respectively. is the non-target signal relative to the qth speech source, the signal power It is used as the variance of the current frequency point, and E{} represents the expectation operation.
[0108] Then, based on the Lagrange multiplier method, the solution of equation (19) can be obtained as:
[0109]
[0110] However, for a noisy and reverberant multi-source scene, the steering vector a q Efficient estimation of (k, l) is challenging (this disclosure does not focus on solving this problem). q(k, l) directly acts on equation (21), transforming the solution of the beamforming model into another solution form:
[0111]
[0112] in, A column vector with 1 at the reference channel position and 0 at the other positions; is the target signal covariance matrix.
[0113] Then the key parameter becomes the non-target signal covariance matrix R in equation (22): q,non (k, l) and the target signal covariance matrix R q,tar (k,l).
[0114] Then, based on the recursive form, R q,non (k,l) and R q,tar (k,l) is estimated as follows:
[0115]
[0116] Among them, α=0.85 is a smoothing parameter, which can be reduced when the noise characteristics change rapidly.
[0117] For Z in the above formula q,non (k,l) can be estimated using the general model as follows:
[0118] Z q,non (k,l)=[1-λ q,mask (k,l)]·Z(k,l) (25)
[0119] in, The masking vector can be estimated by the following directivity gain:
[0120] λ q,mask (k,l)≈G q (k,l)=[G q,1 (k,l) G q,2 (k,l)…G q,M (k,l)] T (26)
[0121]
[0122] Among them, G q,1 (k,l),G q,2 (k,l),...,G q,M (k, l) is the directivity gain of each channel signal, W q,G1 (k,l) and W q,G2(k, l) are the filter weight matrices of the maximum directivity beamformer and the maximum nulling beamformer, respectively. m (k,l) is the dereverberation signal of the mth channel, and Φ{} is a piecewise function, as shown in formula (28).
[0123]
[0124] Maximum directivity filter matrix W q,G1 (k,l) can be estimated according to equations (29) and (30) as follows:
[0125]
[0126]
[0127] in, is the steering vector in the θ direction; is the noise power loaded diagonally, used to prevent the white noise gain from being too low; I M is the unit matrix of dimension M×M.
[0128] The filter weight matrix W of the maximum null beamformer q,G2 (k,l) can be estimated according to formula (31) as follows:
[0129]
[0130] R rec (k,l)=A(k,l)Σ rec (k,l)A H (k,l)+R NN (k,l) (32)
[0131] Σ rec (k,l)=rdiag{diag{Σ(k,l)}} (33)
[0132] Σ(k,l)=pinv{A(k,l)}[R ZZ (k,l)-R NN (k,l)]pinv{A H (k,l)} (34)
[0133] A(k,l)=[a1(k,l)…a q (k,l)…a Q (k,l)] (35)
[0134]
[0135] Among them, the speech source noise covariance matrix It can be estimated by the silence segment. is the steering vector joint matrix, diag{} represents the operation of taking the main diagonal elements, rdiag{} represents the operation of reconstructing the diagonal matrix from the row vectors, and pinv{} represents the operation of finding the pseudo-inverse. Maximum nulling filter matrix The nulling capability of the beamformer is improved by eliminating the correlation between signal components.
[0136] Finally, the enhancement and separation of the speech signal are completed through alternating iteration. The specific steps of alternating iteration are as follows:
[0137] The first step is to initialize the parameters: R NN (k,l), Φ n (k,l),Φ s (k,l);
[0138] The second step is to estimate the steering vector: a1(k,l),a2(k,l),...,a q (k,l),...,a Q (k,l);
[0139] The third step is to iteratively estimate the dereverberation signal according to the estimated steering vector: estimate the Kalman linear prediction dereverberation model according to equations (5) to (11), and calculate the dereverberation signal Z(k, l) according to equation (13);
[0140] Step 4: Estimate the directional gain beamformer weight W according to equations (21) to (36): q (k,l), obtain the unified beamforming weight W(k,l);
[0141] Step 5: Calculate the enhanced signals Y1(k, l), Y2(k, l), ..., Y according to formula (18): Q (k,l) and update the time-varying variance
[0142] Step 6: Update the noise covariance matrix Φ according to equations (14) and (15): n (k, l), the expected signal covariance matrix Φs(k, l), and update the speech plus noise covariance matrix Φ u (k,l);
[0143] Step 6: Φ u Substituting (k, l) back into the third step, the iterative operation can be realized. After a few iterations, a stable denoised, dereverberated and speech-separated speech signal can be obtained.
[0144] In the embodiment of this example, the joint denoising, dereverberation and speech separation method based on directional convolution beamforming disclosed in the present invention uses Kalman filtering to build a dereverberation model of multi-channel linear prediction filtering, which can better estimate the reverberation components in a noisy and reverberant environment and better remove the reverberation components. A directional gain model in a multi-sound source scenario is constructed based on the maximum directional beamformer and the maximum null beamformer, and it is used to estimate the covariance matrix of non-target signals, realizing a real-time MVDR beamforming method with robust denoising performance. The linear prediction dereverberation based on Kalman filtering is associated with the convolution beamformer based on directional gain through time-varying variance weighting, and a convolution beamforming model based on directional gain is derived, and robust denoising, dereverberation and speech separation performance is achieved through an alternating iterative algorithm.
[0145] In the embodiment of this example, compared with the prior art, the beneficial effects of the present disclosure are as follows: First, dereverberation is achieved by combining Kalman filtering with multi-channel linear prediction, which can better achieve dereverberation in a noisy and reverberant environment; second, a maximum directivity beamformer and a maximum nulling beamformer are used to construct a directivity gain, and the noise covariance matrix is estimated in real time to achieve robust denoising and speech separation performance; third, a directional gain convolution beamformer is constructed by combining Kalman filtering and beamforming through time-varying variance weighting, which has robust denoising, dereverberation and speech separation performance.
[0146] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the present disclosure.
[0147] Reference Figure 3 , shows a structural block diagram of a speech enhancement and separation device 200 based on directional convolution beamforming of the present disclosure, the device includes an observation signal construction module 210, an error model establishment module 220, a variance response model establishment module 230, a covariance matrix and time-varying variance estimation module 240 and an iterative processing module 250, wherein:
[0148] An observation signal construction module 210 is used to obtain a multi-channel observation signal according to a microphone array with a preset number of array elements, and to construct a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task;
[0149] The error model building module 220 based on Kalman filtering is used to build a Kalman gain linear prediction error model based on the historical observation signal matrix by using Kalman filtering and multi-channel linear prediction;
[0150] A directivity beamforming model building module 230 is used to solve the directivity gain by using a maximum directivity beamformer and a maximum nulling beamformer, update the noise covariance matrix in real time, and build a minimum variance distortion-free response beamforming model based on the directivity gain;
[0151] A speech and noise covariance matrix and time-varying variance estimation module 240, for estimating the speech and noise covariance matrix based on a minimum variance distortionless response beamformer of directivity gain, and estimating the time-varying variance based on the separated speech source;
[0152] The iterative processing module 250 is used to utilize the time-varying variance to jointly establish the Kalman gain linear prediction error model and the directional gain minimum variance distortion-free response beamforming model, establish a directional convolution beamforming model, and enhance and separate the noisy and reverberant multi-source speech signals through alternating iterations.
[0153] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0154] Optionally, an embodiment of the present disclosure further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0155] The embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0156] Figure 4 800 is a block diagram of an electronic device 800 shown in the present disclosure. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0157] Reference Figure 4 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0158] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0159] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0160] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0161] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0162] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0163] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0164] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0165] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0166] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0167] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0168] Figure 5 19 is a block diagram of a computer-readable storage medium 1900 shown in the present disclosure. For example, the computer-readable storage medium 1900 may be provided as a server.
[0169] Reference Figure 5 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0170] The computer readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0171] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0172] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0173] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, but the present disclosure is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present disclosure, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present disclosure and the claims, all of which are within the protection of the present disclosure.
[0174] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present disclosure can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0175] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0176] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0177] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0179] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks or optical disks.
[0180] The above is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A speech enhancement and separation method based on directional convolution beamforming, characterized in that: The method comprises: According to the microphone array with a preset number of array elements, a multi-channel observation signal is obtained, and a historical observation signal matrix suitable for the multi-channel linear prediction dereverberation task is constructed; According to the historical observation signal matrix, a Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction; The maximum directivity beamformer and the maximum nulling beamformer are used to construct a minimum variance distortion-free response beamforming model based on directivity gain. A minimum variance distortionless response beamformer based on directivity gain estimates the speech and noise covariance matrix and estimates the time-varying variance based on the separated speech sources; The time-varying variance is used to jointly establish the Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model to establish a directional convolution beamforming model, and the enhancement and separation of multi-source speech signals containing noise and reverberation are completed through alternating iterations.
2. The method according to claim 1, characterized in that Obtain multi-channel observation signals and construct a historical observation signal matrix suitable for the multi-channel linear prediction dereverberation task, including: Acquire multi-channel observation signals of a preset number of speech sources according to a microphone array with a preset number of array elements; Performing short-time Fourier transform processing on the multi-channel observation signal; Based on the multi-channel observation signals processed by short-time Fourier transform, a historical observation signal matrix suitable for the multi-channel linear prediction dereverberation task is constructed.
3. The method according to claim 1, characterized in that Using Kalman filtering and multi-channel linear prediction, a Kalman gain linear prediction error model is constructed, including: According to the historical observation signal matrix, based on Kalman filtering, a multi-channel Kalman filter model is established; Based on the multi-channel Kalman filter model, a multi-channel linear prediction model is jointly established to establish a dereverberation construction.
4. The method according to claim 1, characterized in that Using the maximum directivity beamformer and the maximum nulling beamformer, a minimum variance distortion-free response beamforming model based on directivity gain is constructed, including: According to the historical observation signal matrix, a dereverberation signal is estimated using a multi-channel linear prediction model based on Kalman filtering; estimating a denoised and separated speech signal using a directional beamformer based on the dereverberated signal; Based on the Kalman gain linear prediction error model, a noise covariance matrix and an expected signal covariance matrix are established by splitting the speech plus noise covariance matrix; Based on the dereverberation signal, a noise covariance model is generated by constructing a model of a minimum variance distortion-free response beam for the noise covariance matrix; Based on the dereverberation signal, a model of a minimum variance distortion-free response beam is constructed by estimating a target signal covariance and a non-target signal covariance matrix based on a maximum directivity beamformer and a maximum nulling beamformer; Based on the constructed minimum variance distortion-free response beamformer, the separated signals of each speech source are obtained, and the covariance matrix of the speech signal is constructed; Based on the dereverberation signal covariance matrix, the noise covariance matrix is obtained by subtracting the speech signal covariance matrix.
5. The method according to claim 1, characterized in that The Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model are combined, including: According to the dereverberation signal, beamforming weights are calculated by using a time-varying variance-weighted covariance matrix to establish a beamforming model; Based on the Lagrange multiplier method, the beamforming model is solved and an equivalent transformation is performed on the solution of the beamforming model; The covariance matrix of non-target signals and the covariance matrix of target signals in the solution of the equivalently transformed beamforming model are estimated.
6. The method according to claim 5, characterized in that The covariance matrix of the non-target signal and the covariance matrix of the target signal in the solution of the equivalently transformed beamforming model are estimated, including: The speech source noise covariance matrix is estimated according to the silent segment, and the correlation items between signal components are eliminated based on the steering vector joint matrix to establish the maximum null-notch filter matrix; According to the directional steering vector, the maximum directivity filter matrix is established by integrating in the directional space and combining with the diagonal loading matrix; Based on the maximum directivity beamformer and the maximum nulling beamformer, estimating the masking vector by directivity gain; Establishing a general model for estimating non-target signals by estimating the masking vector; Based on the recursive form, the estimation of the non-target signal covariance matrix and the target signal covariance matrix is completed through the estimation general model.
7. The method according to claim 1, characterized in that The enhancement and separation of multi-source speech signals containing noise and reverberation are completed through alternating iterations, including: After initializing the parameters, obtain the steering vector estimate; According to the steering vector, a dereverberation filter is estimated to obtain a dereverberation signal; Obtaining uniform beamforming weights by estimating directivity gain beamforming matrix weights; Based on the unified beamforming weights, the speech plus noise covariance matrix is iteratively updated to complete the iterative estimation processing of denoising and separation signals; Based on the iterative estimation processing results of the dereverberation signal, denoising and separation signal, the enhancement and separation of the speech signal are achieved.
8. A speech enhancement and separation device based on directional convolution beamforming, characterized in that: The device comprises: An observation signal construction module is used to obtain multi-channel observation signals according to a microphone array with a preset number of array elements, and to construct a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task; An error model building module is used to build a Kalman gain linear prediction error model based on the historical observation signal matrix using Kalman filtering and multi-channel linear prediction; A variance response model building module is used to build a minimum variance distortion-free response beamforming model based on directivity gain by using a maximum directivity beamformer and a maximum nulling beamformer; A covariance matrix and time-varying variance estimation module, for estimating speech and noise covariance matrices based on a directional minimum variance distortionless response beamformer, and estimating time-varying variance based on separated speech sources; The iterative processing module is used to utilize the time-varying variance to jointly establish the Kalman gain linear prediction error model and the minimum variance distortion-free response beamforming model, establish a directional convolution beamforming model, and enhance and separate the noisy and reverberant multi-source speech signals through an alternating iterative method.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and system for dereverberation based on Kalman filtering
CN108172231A
Generation method and system for adaptive beam former
CN110687528A
Microphone array beam forming method
CN110931036A
Microphone array voice beam forming system based on OMAP-L137
CN113113037A
Sound source signal extraction method and system based on geometric constraint source extraction and dereverberation
CN117334213A