Speech enhancement and separation method and device based on directional convolution beamforming
A speech signal enhancement and separation model was constructed by combining directional convolutional beamforming with Kalman filtering and multi-channel linear prediction. Through alternating iteration, speech signal enhancement and separation were achieved.
Patent Information
- Application Number
- CN202411981375.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing convolutional beamformers are insufficient in removing early reverberation and time-varying variance sensitivity, which affects the clarity and naturalness of speech signals. Furthermore, multi-channel linear prediction methods cannot achieve optimal results when jointly optimized.
A directional convolutional beamforming method is adopted, which combines Kalman filtering and multi-channel linear prediction to construct a Kalman gain linear prediction error model. Then, using a maximum directional beamformer and a maximum null beamformer, speech enhancement and speech enhancement and separation are performed in an alternating iterative manner.
It achieves the enhancement and separation of speech signals.
Smart Images

Figure CN119943085B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of acoustic signal processing, and in particular to a speech enhancement and separation method and device based on directional convolution beamforming. BACKGROUND
[0002] In microphone array-based speech enhancement, beamforming technology is one of the commonly used filtering methods, and its main goal is to improve signal quality by spatial filtering technology according to the spatial distribution of acoustic components such as sound sources, noise and reverberation, to achieve denoising, dereverberation and sound source separation. Classic beamformers, such as delay-and-sum beamformer, minimum power distortion response (MPDR) beamformer, minimum variance distortion response (MVDR) beamformer, etc., although perform well in noise suppression, but perform weakly in dereverberation. This problem has triggered a large number of studies based on multi-channel linear prediction, especially the weighted power error (WPE) algorithm. The WPE algorithm achieves effective dereverberation performance by modeling the linear prediction relationship of multi-channel time series signals; therefore, in the past, many studies have cascaded the WPE algorithm with the beamformer to achieve better denoising and dereverberation performance. However, WPE and beamformer usually work independently, which leads to suboptimal performance when jointly optimized.
[0003] In the prior art, multi-channel linear prediction methods based on convolution beamformers (CBF) have been proposed, such as weighted power minimum distortion response (WPD) and source-accurate decomposition-based convolution beamformer (SWF-CBF). These methods combine beamforming and multi-channel linear prediction for dereverberation by jointly optimizing the cost function, aiming to achieve better denoising and dereverberation performance. Although this technology based on convolution beamformers has made some progress, there are still some problems that cannot be ignored. First, multi-channel linear prediction methods are usually only suitable for late reverberation removal, which leads to residual early reverberation components in the separated sound source signals in actual applications, affecting the clarity and naturalness of speech. Second, many convolution beamformers are based on MPDR beamformers, and the performance of convolution beamformers based on MPDR is often dependent on the accuracy of time-varying variance. However, in actual environments, the estimation of time-varying variance often has errors, which can lead to distortion of the target sound source and residual of the interfering sound source.
[0004] Therefore, although the joint optimization of convolution beamformers can improve the dereverberation performance to some extent, it still faces great challenges in reducing early reverberation and being sensitive to time-varying variance, and effective solutions are urgently needed. SUMMARY
[0005] The present disclosure shows a speech enhancement and separation method and device based on directional convolution beamforming.
[0006] In a first aspect, the disclosure shows a voice enhancement and separation method based on directivity convolution beamforming, the method comprising:
[0007] According to the microphone array with a preset number of array elements, a multi-channel observation signal is obtained, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed;
[0008] According to the history observation signal matrix, a Kalman gain linear prediction error model is constructed using Kalman filtering and multi-channel linear prediction;
[0009] A minimum variance distortionless response beamforming model based on directivity gain is constructed using a maximum directivity beamformer and a maximum null beamformer;
[0010] The minimum variance distortionless response beamforming model based on directivity gain estimates the speech plus noise covariance matrix, and estimates the time-varying variance based on the separated speech sources;
[0011] The time-varying variance is used to simultaneously solve the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model, a directivity convolution beamforming model is established, and the enhancement and separation of the multi-source speech signal containing noise and reverberation are completed through an alternating iteration method.
[0012] In an exemplary embodiment of the disclosure, a multi-channel observation signal is obtained, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed, comprising:
[0013] According to the microphone array with a preset number of array elements, a multi-channel observation signal of a preset number of speech sources is obtained;
[0014] The multi-channel observation signal is subjected to short-time Fourier transform processing;
[0015] Based on the multi-channel observation signal after short-time Fourier transform processing, a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed.
[0016] In an exemplary embodiment of the disclosure, a Kalman gain linear prediction error model is constructed using Kalman filtering and multi-channel linear prediction, comprising:
[0017] Based on the history observation signal matrix, a multi-channel Kalman filtering model is established based on Kalman filtering;
[0018] Based on the multi-channel Kalman filtering model, a dereverberation construction is established by simultaneously solving a multi-channel linear prediction model.
[0019] In an exemplary embodiment of the present disclosure, a minimum variance distortionless response beamforming model based on directivity gain is constructed using a maximum directivity beamformer and a maximum null beamformer, comprising:
[0020] According to the historical observation signal matrix, a de-reverberation signal is estimated using a multi-channel linear prediction model based on Kalman filtering;
[0021] According to the de-reverberation signal, a de-noised and separated speech signal is estimated using a directivity beamformer;
[0022] Based on the Kalman gain linear prediction error model, a noise covariance matrix and an expected signal covariance matrix are established by splitting the speech plus noise covariance matrix;
[0023] Based on the de-reverberation signal, a target signal covariance matrix and a non-target signal covariance matrix are estimated based on a maximum directivity beamformer and a maximum null beamformer, and a minimum variance distortionless response beamforming model is constructed;
[0024] Based on the constructed minimum variance distortionless response beamforming model, a separated signal of each speech source is obtained, and a covariance matrix of the speech signal is constructed;
[0025] Based on the de-reverberation signal covariance matrix, a noise covariance matrix is obtained by subtracting the speech signal covariance matrix.
[0026] In an exemplary embodiment of the present disclosure, the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model are solved simultaneously, comprising:
[0027] According to the de-reverberation signal, the beamforming weight is calculated by a time-varying variance weighted covariance matrix, and a convolution beamforming model is established;
[0028] Based on the Lagrange multiplier method, the minimum variance beamforming model is solved, and the solution of the beamforming model is equivalently transformed;
[0029] The non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalently transformed beamforming model are estimated.
[0030] In an exemplary embodiment of the present disclosure, the non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalently transformed beamforming model are estimated, comprising:
[0031] Based on the noise covariance matrix estimated in the silence segment, and based on the steering vector joint matrix, the correlation terms between signal components are eliminated, and a maximum null beamformer is established;
[0032] According to the direction orientation vector, integration is performed in the direction space, and a maximum directivity beamformer is established in combination with a diagonal loading matrix;
[0033] Based on the maximum directivity beamformer and the maximum null beamformer, a masking vector is estimated through a directivity gain;
[0034] An estimated general model of a non-target signal is established through the estimated masking vector;
[0035] Based on a recursive form, estimation of the non-target signal covariance matrix and the target signal covariance matrix is completed through the estimated general model.
[0036] In an exemplary embodiment of the present disclosure, enhancement and separation of a multi-source speech signal containing noise and reverberation are completed through an alternating iterative manner, comprising:
[0037] Initialization parameters are estimated, and a steering vector is estimated;
[0038] According to the estimated steering vector, a dereverberation filter is estimated, and a dereverberation signal is obtained;
[0039] Through estimation of a directivity gain beamforming matrix weight, a unified beamforming weight is obtained;
[0040] Based on the unified beamforming weight, the speech noise covariance matrix and the time-varying variance are iteratively updated;
[0041] Through alternating iteration of a multi-channel linear prediction based on Kalman filtering and a directivity gain based beamformer, the performance of convolution beamforming is realized, and enhancement and separation of a speech signal are completed.
[0042] In a second aspect, the present disclosure shows a speech enhancement and separation device based on directivity convolution beamforming, comprising:
[0043] An observation signal construction module is configured to obtain a multi-channel observation signal from a microphone array with a preset number of array elements, and construct a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task;
[0044] An error modeling module based on Kalman filtering is configured to construct a Kalman gain linear prediction error model by using Kalman filtering and multi-channel linear prediction based on the historical observation signal matrix;
[0045] A directivity beamforming modeling module is configured to construct a minimum mean square error response beamforming model based on a directivity gain by combining a maximum directivity beamformer and a maximum null beamformer;
[0046] a speech-plus-noise covariance matrix and time-varying variance estimation module configured to estimate a speech-plus-noise covariance matrix based on a directivity gain minimum variance distortion response beamforming model and estimate a time-varying variance based on separated speech sources;
[0047] an alternating iteration processing module configured to use the time-varying variance to simultaneously solve the Kalman gain linear prediction error model, the directivity gain minimum variance distortion response beamforming model, establish a directivity convolution beamforming model, and complete enhancement and separation of the noisy and unclear response multi-speech source speech signal through alternating iteration.
[0048] In a third aspect, the present disclosure shows an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method of any one of the above aspects.
[0049] In a fourth aspect, the present disclosure shows a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the method of any one of the above aspects.
[0050] In a fifth aspect, the present disclosure shows a computer program product, when the instructions in the computer program product are executed by the processor of an electronic device, the electronic device can execute the method of any one of the above aspects.
[0051] The speech enhancement and separation method based on directivity convolution beamforming of the present disclosure first constructs a historical observation signal matrix from a multi-channel observation signal obtained according to a microphone array with a preset number of array elements; secondly, a Kalman gain linear prediction error model based on Kalman filtering and multi-channel linear prediction de-reverberation is constructed according to the historical observation signal matrix; then, a directivity gain minimum variance distortion response beamforming model is constructed based on a maximum directivity beamformer and a maximum null beamformer; next, a speech-plus-noise covariance matrix is estimated based on the directivity gain minimum variance distortion response beamforming model, and a time-varying variance is estimated based on separated speech sources; finally, the Kalman gain linear prediction error model based on Kalman filtering and the directivity gain minimum variance distortion response beamforming model are simultaneously solved using the time-varying variance, a directivity convolution beamforming model is established, and enhancement and separation of the speech signal are completed through alternating iteration.
[0052] The present disclosure utilizes Kalman filtering combined with multi-channel linear prediction to realize dereverberation, which can better realize dereverberation in a noisy and reverberant environment. At the same time, a directivity gain is constructed by using a maximum directivity beamformer and a maximum null beamformer, and a noise covariance matrix is estimated in real time to achieve robust denoising and speech separation performance. In addition, further, by time-varying variance weighting, the multi-channel linear prediction model based on Kalman filtering and the minimum variance distortionless response beamforming model based on directivity gain are combined to construct a directivity gain convolution beamformer, which has robust denoising, dereverberation and speech separation performance. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a step flow chart of a speech enhancement and separation method based on directivity convolution beamforming of the present disclosure.
[0054] Figure 2 is an algorithm principle diagram of a speech enhancement and separation method based on directivity convolution beamforming of the present disclosure.
[0055] Figure 3 is a structure block diagram of a speech enhancement and separation device based on directivity convolution beamforming of the present disclosure.
[0056] Figure 4 is a block diagram of an electronic device of the present disclosure.
[0057] Figure 5 is a block diagram of a computer readable storage medium of the present disclosure. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.
[0059]
NOUN EXPLANATION
[0060] The convolutional beamformer (CBF) refers to a spatial filter constructed by using multi-channel linear prediction and beamforming based on multi-channel observation signals at past time, which is commonly used in speech signal denoising, dereverberation and interference sound source suppression.
[0061] The minimum variance distortionless response (MVDR) beamformer is an adaptive beamformer which is constructed by constraining the target direction distortionless while minimizing the non-target signal output power, and is commonly used in the field of signal enhancement processing such as sonar, radar, communication, speech, etc.
[0062] The multi-channel linear prediction (MCLP) refers to a technique of estimating the true value signal at the current time in the form of linear prediction filtering by using the multi-channel observation signals at the past time, and is commonly used in the late reverberation suppression of the speech signal.
[0063] Referring to Figure 1 , a step flow chart of a speech enhancement and separation method based on directional convolution beamforming of the present disclosure is shown, and the method can be applied to an electronic device, wherein the method can specifically include the following steps:
[0064] In step S110, a multi-channel observation signal is obtained according to a microphone array with a preset number of array elements, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed;
[0065] In step S120, a Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction according to the history observation signal matrix;
[0066] In step S130, a minimum variance distortionless response beamforming model based on directional gain is constructed by using a maximum directivity beamformer and a maximum null beamformer;
[0067] In step S140, a speech plus noise covariance matrix is estimated based on the minimum variance distortionless response beamforming model based on directional gain, and a time-varying variance is estimated based on the separated speech source;
[0068] In step S150, the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model based on directional gain are solved by using the time-varying variance, a directional convolution beamforming model is established, and the enhancement and separation of the multi-source speech signal containing noise and reverberation are completed by an alternating iteration method.
[0069] The speech enhancement and separation method based on directivity convolution beamforming of the present disclosure utilizes Kalman filtering combined with multi-channel linear prediction to more effectively remove reverberation and improve dereverberation performance in a noisy reverberation environment. Secondly, based on the directivity gain model constructed by the maximum directivity beamformer and the maximum null beamformer, the noise covariance matrix can be estimated in real time to achieve robust denoising and speech separation. In addition, by combining the multi-channel linear prediction based on Kalman filtering with the minimum variance distortionless response beamforming technology based on directivity gain through time-varying variance weighting, a directivity gain convolution beamformer is constructed to improve the overall robustness of denoising, dereverberation and speech separation. These advantages enable the method to have stronger processing capability in a complex noise and reverberation environment.
[0070] In the embodiment of the present example, as shown in Figure 2 The joint denoising, dereverberation and interference speech suppression method based on the directivity convolution beamformer of the present disclosure is implemented by the following technical solutions:
[0071] In step S110, according to the microphone array with a preset number of elements, a multi-channel observation signal is obtained, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed.
[0072] Firstly, let the number of elements of the microphone array be M , and the number of speech sources be Q , then the multi-channel observation signal x(t) of the preset speech source can be obtained, as shown in equation (1):
[0073] (1)
[0074] wherein, q =1,2,..., Q is the speech source index, is the observation signal collected by the microphone array at the historical t time, is the direct sound signal of the q th speech source, is the reflected sound signal (reverberation) of the q th speech source, is the noise collected by the microphone array at the historical t time, represents real-time space.
[0075] Then, short-time Fourier transform (STFT) is performed on equation (1) to obtain equation (2):
[0076] (2)
[0077] wherein, are respectively x(t ), s q ( t ), r q ( t ) and n t ) are the STFT results of s k , r l and n , respectively, are the indices of frequency and frame number,
[0078] Based on the STFT processed formula (2), a history observation signal matrix suitable for multi-channel linear prediction dereverberation task is constructed:
[0079] (3)
[0080] wherein, is the stacked signal of the history observation signal, b is the linear prediction delay, L g is the order of the linear prediction filter, and the superscript “ T ” represents the transposition operation.
[0081] In step S120, a Kalman gain linear prediction error model is constructed based on the history observation signal matrix by using Kalman filtering and multi-channel linear prediction.
[0082] Firstly, the multi-channel linear prediction filter is set to wherein, L G ML g The Kalman gain linear prediction error model can be modeled as follows:
[0083] (4)
[0084] wherein, is the estimated Kalman linear prediction dereverberation model, E{} represents the expectation operation, and the superscript “ H ” represents the conjugate transposition operation.
[0085] However, since G k , l is unknown in practice, the Kalman filtering is used to estimate and calculate , as shown in formulas (5)-(11):
[0086] (5)
[0087] (6)
[0088] (7)
[0089] (8)
[0090] (9)
[0091] (10)
[0092] (11)
[0093] wherein, is a priori estimate of , and are identity matrices of dimension and respectively, is the difference between the observed signal and the linearly predicted reverberation component, is a sparse matrix of , denotes the Kronecker product, is a multi-channel Kalman gain, is a process noise covariance matrix, is an estimated speech-plus-noise covariance matrix.
[0094] In step S130, a minimum variance distortionless response beamforming model based on directivity gain is constructed by using a maximum directivity beamformer and a maximum null distortionless response beamformer.
[0095] In step S140, the speech-plus-noise covariance matrix is estimated based on the minimum variance distortionless response beamforming model based on directivity gain, and the time-varying variance is estimated based on the separated speech sources.
[0096] The Kalman gain K( k , l ) shown in formula (9) is a key parameter in Kalman filtering. In single-channel Kalman filtering, the Kalman gain can be regarded as: error signal x filter error / (error signal power + expected signal power), which is used to correct the filtering error of the current frequency point; and in multi-channel Kalman filtering, the Kalman gain becomes: error signal vector x filter error / (error signal covariance matrix + expected signal covariance matrix), which is used to correct the filtering error of the current frequency vector. In single-channel processing, in order to improve the accuracy of the Kalman gain, many studies focus on the estimation of the expected signal power. In the present disclosure, the speech-plus-noise covariance matrix is split into two parts: the noise covariance matrix and the expected signal covariance matrix by the minimum variance distortionless response beamforming model, and a certain method is used for estimation respectively.
[0097] The first step, let the dereverberated signal be , based on can be estimated as follows:
[0098] (12)
[0099] The second step, let the number of speech sources be Q ( Q M ), then the beamforming weight of the q th speech source is , q =1,2,..., Q , then the signal after de-noising, dereverberation and separation can be estimated as follows:
[0100] (13)
[0101] The third step, based on the formula (12) and (13), the noise covariance matrix Φ n ( k , l ) and the expected signal covariance matrix Φ s ( k , l ) can be estimated as follows:
[0102] (14)
[0103] (15)
[0104] (16)
[0105] wherein, P is the estimated power of the q th speech source.
[0106] The fourth step, since each speech source is separated one by one according to the formula (13) and needs Q beamformers to process, in order to make the beamformers separate each speech source uniformly, the joint beamforming weight matrix is represented as follows:
[0107] (17)
[0108] wherein, the uniformly separated signal can be represented as:
[0109] (18)
[0110] wherein, disg{} represents taking the main diagonal elements.
[0111] Fifth step, substitute equation (18) into equations (14) and (15) to obtain the noise covariance matrix Φ. n ( k , l The estimation results of the expected signal covariance matrix Φs(k,l).
[0112] In step S150, the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model are combined using time-varying variance to establish a directional convolution beamforming model, and the enhancement and separation of noisy and reverberant multi-source speech signals are completed through alternating iteration.
[0113] First, based on the dereverberation signal Z( k , l A beamforming model is established by calculating the beamforming weights using a time-varying variance-weighted covariance matrix.
[0114] (19)
[0115] (20)
[0116] in, The first q Time-varying variance-weighted covariance matrix and steering vector of the non-target signal of each speech source For relative to the first q The non-target signal from the speech source, signal power It is used as the variance of the current frequency point, and E{} represents the expectation operation.
[0117] Then, based on the Lagrange multiplier method, the solution to equation (19) can be obtained as follows:
[0118] (twenty one)
[0119] However, for noisy, reverberant multi-source scenes, the guiding vector a q ( k , l Efficient estimation of ) is challenging (this disclosure does not focus on solving this problem). To avoid inaccurate a q ( k , l Directly applied to equation (21), the solution to the beamforming model is equivalently transformed into another solution form:
[0120] (twenty two)
[0121] in, A column vector with 1 for the reference channel position and 0 for the rest; Let be the covariance matrix of the target signal.
[0122] The key parameter then becomes the non-target signal covariance matrix R in equation (22). q,non ( k , l ) and the target signal covariance matrix R q,tar ( k , l ).
[0123] Then, based on the recursive form, R can be... q,non ( k , l ) and R q,tar ( k , l The estimate is as follows:
[0124] (twenty three)
[0125] (twenty four)
[0126] in, α =0.85 is a smoothing parameter, which can be reduced when the noise characteristics change rapidly.
[0127] For Z in the above formula q,non ( k , l The general model can be represented as follows:
[0128] (25)
[0129] in, The masking vector can be estimated using the following directional gain:
[0130] (26)
[0131] (27)
[0132] in, For the directional gain of each channel signal, These are the filter weight matrices for the maximum directivity beamformer and the maximum null beamformer, respectively. No. m De-revering signal for each channel It is a piecewise function, as shown in equation (28).
[0133] (28)
[0134] Maximum Directivity Filter Matrix The can be estimated as follows according to the formula (29), (30):
[0135] (29)
[0136] (30)
[0137] wherein, the steering vector of the direction; is the diagonal loaded noise power, used to prevent the white noise gain from being too low; M is the unit matrix of dimension M x M.
[0138] The filter weight matrix W of the maximum null beamformer q,G2 k l The can be estimated as follows according to the formula (31):
[0139] (31)
[0140] (32)
[0141] (33)
[0142] (34)
[0143] (35)
[0144] (36)
[0145] wherein, is the noise covariance matrix estimated by the mute section, is the steering vector joint matrix, diag{} represents the operation of taking the main diagonal elements, rdiag{} represents the operation of reconstructing the diagonal matrix from the row vector, pinv{} represents the operation of finding the pseudo-inverse. The maximum null filter matrix The nulling ability of the beamformer is improved by eliminating the correlation terms between the signal components.
[0146] Finally, the enhancement and separation of the speech signal are completed by the alternating iteration. The specific implementation steps of the alternating iteration are as follows:
[0147] Firstly, the parameters are initialized: NN k l n k l s k , l );
[0148] Second step, estimate steering vectors: a1( k , l ), a2( k , l ),..., a q ( k , l ),..., a Q ( k , l );
[0149] Third step, according to the estimated steering vectors, iteratively estimate the dereverberated signal: estimate Kalman linear prediction dereverberation model according to formula (5)-(11), calculate dereverberated signal Z( k , l ) according to formula (13);
[0150] Fourth step, estimate directivity gain beamformer weight W q ( k , l ) according to formula (21)-(36), obtain unified beamforming weight W( k , l );
[0151] Fifth step, calculate enhanced signal Y 1( k , l ), Y 2( k , l ),..., Y Q ( k , l ) according to formula (18) and update time-varying variance ;
[0152] Sixth step, update noise covariance matrix Φ n ( k , l ), expected signal covariance matrix Φs( k , l ) according to formula (14) and formula (15), and update speech plus noise covariance matrix Φ u ( k , l );
[0153] Sixth step, Φ u ( k , l) The third step is returned to, and the iterative operation can be realized. After a few iterations, the stable de-noising, dereverberation and speech separation speech signal can be obtained.
[0154] In the embodiment of the present example, the joint de-noising, dereverberation and speech separation method based on the directional convolution beamforming of the present disclosure utilizes the Kalman filter to build a dereverberation model of the multi-channel linear prediction filter, and can better estimate and remove the reverberation component in the noisy and reverberant environment. The directional gain model in the multi-source scene is constructed based on the maximum directivity beamformer and the maximum null beamformer, and is used for estimating the non-target signal covariance matrix, so as to realize the real-time MVDR beamforming method with robust de-noising performance. The linear prediction dereverberation based on the Kalman filter and the convolution beamformer based on the directional gain are associated through the time-varying variance weighting, the convolution beamforming model based on the directional gain is derived, and the robust de-noising, dereverberation and speech separation performance is realized through the alternating iterative algorithm.
[0155] In the embodiment of the present example, compared with the prior art, the present disclosure has the following beneficial effects: first, the Kalman filter is used to realize the dereverberation through the multi-channel linear prediction, so that the dereverberation can be better realized in the noisy and reverberant environment; second, the maximum directivity beamformer and the maximum null beamformer are used to construct the directional gain, so that the noise covariance matrix can be estimated in real time, and the robust de-noising and speech separation performance can be achieved; third, the Kalman filter and the beamforming are combined through the time-varying variance weighting, the directional gain convolution beamformer is constructed, and the robust de-noising, dereverberation and speech separation performance is achieved.
[0156] It should be noted that, for the method embodiment, in order to simply describe, it is expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the action sequence described, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily essential to the present disclosure.
[0157] Referring to Figure 3 , a structure block diagram of a speech enhancement and separation device 200 based on directional convolution beamforming of the present disclosure is shown, the device comprises an observation signal construction module 210, an error model establishment module 220, a variance response model establishment module 230, a covariance matrix and time-varying variance estimation module 240 and an iterative processing module 250, wherein:
[0158] The observation signal construction module 210 is used for acquiring a multi-channel observation signal according to a microphone array with a preset number of array elements, and constructing a historical observation signal matrix suitable for a multi-channel linear prediction dereverberation task.
[0159] The Kalman filter-based error model establishing module 220 is configured to construct a Kalman gain linear prediction error model by using Kalman filtering and multi-channel linear prediction according to the historical observation signal matrix.
[0160] The directivity beamforming model establishing module 230 is configured to solve a directivity gain, update a noise covariance matrix in real time, and construct a minimum variance distortion response beamforming model based on the directivity gain by using a maximum directivity beamformer and a maximum null beamformer.
[0161] The speech plus noise covariance matrix and time-varying variance estimation module 240 is configured to estimate a speech plus noise covariance matrix based on the minimum variance distortion response beamforming model based on the directivity gain, and estimate a time-varying variance based on a separated speech source.
[0162] The iterative processing module 250 is configured to solve the Kalman gain linear prediction error model and the directivity gain minimum variance distortion response beamforming model by using the time-varying variance, establish a directivity convolution beamforming model, and complete enhancement and separation of a multi-speech source speech signal containing noise and reverberation by using an alternating iteration method.
[0163] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts are referred to the part of the method embodiment.
[0164] Optionally, the embodiment of the disclosure further provides an electronic device, which comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program is executed by the processor to implement each process of the method embodiment and achieve the same technical effects, and details are not repeated here.
[0165] The embodiment of the disclosure further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, wherein the computer program is executed by a processor to implement each process of the method embodiment and achieve the same technical effects, and details are not repeated here. The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0166] Figure 4 is a block diagram of an electronic device 800 shown in the disclosure. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a message transmission device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0167] Referring to Figure 4 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0168] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the above methods. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0169] The memory 804 is configured to store various types of data to support operations of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0170] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 can include a power supply management system, one or more power supplies, and other components associated with generating, managing and distributing power for the electronic device 800.
[0171] The multimedia component 808 includes a screen to provide an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, or a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. The front camera and / or the rear camera can receive external multimedia data when the electronic device 800 is in an operating mode, such as a shooting mode or a video mode. Each of the front and rear camera can be a fixed optical lens system or have a focal length and optical zooming capability.
[0172] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) to receive an external audio signal when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker to output an audio signal.
[0173] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0174] The sensor component 814 includes one or more sensors to provide various state assessments for the electronic device 800. For example, the sensor component 814 can detect an open / closed state of the device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, an orientation or acceleration / deceleration of the electronic device 800, and a temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 814 can further include a light sensor such as a CMOS or CCD image sensor for use in an imaging application. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0175] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a cellular network (e.g., 2G, 3G, 4G or 5G), or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast operation information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques and other techniques.
[0176] In an example embodiment, the electronic device 800 can be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic elements, for performing the above-described methods.
[0177] In an example embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to complete the above-described methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0178] Figure 5 is a block diagram of a computer-readable storage medium 1900 according to an example embodiment of the present disclosure. For example, the computer-readable storage medium 1900 can be provided as a server.
[0179] Referring to Figure 5 The computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described methods.
[0180] The computer-readable storage medium 1900 can also include a power supply component 1926 configured to perform power management for the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0181] It should be noted that, in the present document, the terms "comprising", "containing" or any other similar term are intended to encompass non-exclusive inclusions, such that a process, a method, an article or an apparatus that comprises a list of elements does not only include those elements, but can also include other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes that element.
[0182] From the above description of the embodiments, it can be clear to those skilled in the art that the above-mentioned example methods can be implemented by means of software plus a general purpose hardware platform, of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) and includes a number of instructions for making a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) execute the methods described in various embodiments of the present disclosure.
[0183] The embodiments of the present disclosure are described above in combination with the accompanying drawings, but the present disclosure is not limited to the above-described specific embodiments, which are merely illustrative rather than limiting, and those of ordinary skill in the art can make many forms without departing from the purpose of the present disclosure and the scope protected by the claims under the inspiration of the present disclosure, which all belong to the protection of the present disclosure.
[0184] Those skilled in the art can clearly understand the unit and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present disclosure can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0185] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0186] In the embodiments provided by the present disclosure, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is merely a logical function division, and another division manner can be used in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0187] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0188] In addition, each functional unit in each embodiment of the present disclosure can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0189] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure, essentially or in part, or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as U disk, mobile hard disk, ROM, RAM, magnetic disk, or optical disk.
[0190] The above is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for speech enhancement and separation based on directional convolution beamforming, characterized in that, The method comprises: According to the microphone array of the preset array element number, the multi-channel observation signal is obtained, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed; According to the history observation signal matrix, a Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction; A minimum variance distortionless response beamforming model based on directivity gain is constructed by using a maximum directivity beamformer and a maximum null beamformer; A speech plus noise covariance matrix is estimated based on the minimum variance distortionless response beamforming model based on directivity gain, and a time-varying variance is estimated based on a separated speech source; The Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model are solved by using the time-varying variance, a directivity convolution beamforming model is established, and the enhancement and separation of a multi-source speech signal containing noise and reverberation are completed by an alternating iteration method; The Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model are solved, comprising: According to the dereverberation signal, the beamforming weight is calculated by a time-varying variance weighted covariance matrix, and a beamforming model is established; Based on the Lagrange multiplier method, the beamforming model is solved, and an equivalent transformation is performed on the solution of the beamforming model; The non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalent transformed beamforming model are estimated; The non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalent transformed beamforming model are estimated, comprising: According to the noise covariance matrix estimated based on the silence segment, and based on a steering vector joint matrix, a maximum null filter matrix is established by eliminating the correlation terms between signal components; According to the directional steering vector, integration is performed in the directional space, and a maximum directivity filter matrix is established by combining a diagonal loading matrix; Based on the maximum directivity beamformer and the maximum null beamformer, a masking vector is estimated by directivity gain; An estimated general model of the non-target signal is established by the estimated masking vector; Based on the recursive form, the estimated general model is used to complete the estimation of the non-target signal covariance matrix and the target signal covariance matrix.
2. The method of claim 1, wherein, The multi-channel observation signal is obtained, and a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed, comprising: According to the microphone array of the preset array element number, a preset number of multi-channel observation signals of speech sources are obtained; The multi-channel observation signals are processed by short-time Fourier transform; Based on the multi-channel observation signals processed by short-time Fourier transform, a history observation signal matrix suitable for a multi-channel linear prediction dereverberation task is constructed.
3. The method of claim 1, wherein, The Kalman gain linear prediction error model is constructed by using Kalman filtering and multi-channel linear prediction, comprising: According to the history observation signal matrix, a multi-channel Kalman filtering model is established based on Kalman filtering; Based on the multi-channel Kalman filtering model, a dereverberation construction is established by combining a multi-channel linear prediction model.
4. The method of claim 1, wherein, The minimum variance distortionless response beam forming model based on directivity gain is constructed by using the maximum directivity beam former and the maximum null beam former, including: According to the history observation signal matrix, the dereverberation signal is estimated by using the multi-channel linear prediction model based on Kalman filtering; According to the dereverberation signal, the noise-removed and separated speech signal is estimated by using the directivity beam former; Based on the Kalman gain linear prediction error model, the noise covariance matrix and the expected signal covariance matrix are established by splitting the speech plus noise covariance matrix; Based on the dereverberation signal, the noise covariance model is generated by constructing the minimum variance distortionless response beam model of the noise covariance matrix; Based on the dereverberation signal, the target signal covariance and the non-target signal covariance matrix are estimated by using the maximum directivity beam former and the maximum null beam former, and the minimum variance distortionless response beam model is constructed; Based on the constructed minimum variance distortionless response beam forming model, the separated signals of each speech source are obtained, and the covariance matrix of the speech signal is constructed; Based on the dereverberation signal covariance matrix, the noise covariance matrix is obtained by subtracting the speech signal covariance matrix.
5. The method of claim 1, wherein, The enhancement and separation of the multi-source speech signal containing noise and reverberation are completed by alternating iteration, including: After initializing the parameters, the steering vector estimation is obtained; According to the steering vector, the dereverberation filter is estimated, and the dereverberation signal is obtained; By estimating the directivity gain beam forming matrix weight, the unified beam forming weight is obtained; Based on the unified beam forming weight, the speech plus noise covariance matrix is iteratively updated, and the iterative estimation processing of the noise-removed and separated signal is completed; Based on the iterative estimation processing result of the dereverberation signal and the noise-removed and separated signal, the enhancement and separation of the speech signal are realized.
6. A speech enhancement and separation apparatus based on directional convolution beamforming, characterized by The device includes: The observation signal construction module is used for obtaining multi-channel observation signals from a microphone array with a preset number of array elements, and constructing a history observation signal matrix suitable for multi-channel linear prediction dereverberation tasks; The error model establishment module is used for constructing a Kalman gain linear prediction error model by using Kalman filtering and multi-channel linear prediction according to the history observation signal matrix; The variance response model establishment module is used for constructing a minimum variance distortionless response beam forming model based on directivity gain by using a maximum directivity beam former and a maximum null beam former; The covariance matrix and time-varying variance estimation module is used for estimating a speech plus noise covariance matrix based on the minimum variance distortionless response beam forming model based on directivity gain, and estimating a time-varying variance based on the separated speech sources; The iterative processing module is used for using the time-varying variance to simultaneously solve the Kalman gain linear prediction error model and the minimum variance distortionless response beam forming model, establishing a directivity convolution beam forming model, and completing the enhancement and separation of the multi-source speech signal containing noise and reverberation by alternating iteration; The simultaneous solution of the Kalman gain linear prediction error model and the minimum variance distortionless response beam forming model includes: According to the dereverberation signal, a beamforming weight is calculated by a time-varying variance weighted covariance matrix, and a beamforming model is established; Based on the Lagrange multiplier method, the beamforming model is solved, and an equivalent transformation is performed on the solution of the beamforming model; The non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalent transformed beamforming model are estimated; The non-target signal covariance matrix and the target signal covariance matrix in the solution of the equivalent transformed beamforming model are estimated, including: According to the noise covariance matrix estimated based on the mute section, and based on a steering vector joint matrix, a signal component correlation term is eliminated, and a maximum null filter matrix is established; According to the directional steering vector, integration is performed in the directional space, and a maximum directivity filter matrix is established by combining a diagonal loading matrix; Based on the maximum directivity beamformer and the maximum null beamformer, a masking vector is estimated by a directivity gain; An estimated general model of a non-target signal is established by the estimated masking vector; Based on a recursive form, the non-target signal covariance matrix and the target signal covariance matrix are estimated by the estimated general model.
7. An electronic device, comprising: Including: A processor, a memory, and a computer program stored on the memory and executable on the processor, when the computer program is executed by the processor, the method as claimed in any one of claims 1 to 5 is realized.
8. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and when the computer program is executed by the processor, the method as claimed in any one of claims 1 to 5 is realized.
Citation Information
Patent Citations
Method and system for dereverberation based on Kalman filtering
CN108172231A
Generation method and system for adaptive beam former
CN110687528A