A generalized sidelobe elimination method for microphone arrays
Through the sparse blocking matrix and multi-frame expansion method, combined with fixed beam formation and adaptive interference cancellation filter, the problem of high computing power consumption in microphone array noise suppression and speech enhancement is solved, and the effect of generalized sidelobe cancellation of microphone array is improved, and efficient microphone array noise suppression and speech enhancement is achieved.
Patent Information
- Application Number
- CN202211393142.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-11-08
AI Technical Summary
The existing microphone array noise suppression and voice enhancement methods consume too much computing power during multi-frame data processing, making it difficult to efficiently realize the generalized sidelobe elimination of microphone arrays in real-time processing systems.
The sparse blocking matrix construction and multi-frame expansion method is adopted to reduce the data dimension by constructing the sparse blocking matrix, utilize the correlation between multi-frame data, and combine a fixed beamforming filter and an adaptive interference cancellation filter to reduce computing power consumption and improve the generalized sidelobe elimination effect of microphone array.
While reducing computing power consumption, the performance of the generalized sidelobe elimination algorithm of microphone array is improved, a large amount of computing resources is saved, and the noise reduction and voice enhancement effect of microphone array is improved.
Smart Images

Figure CN115691530B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of speech processing technology, in particular to the field of microphone array noise suppression and speech enhancement, and specifically relates to a microphone array generalized sidelobe elimination method. Background Art
[0002] Common microphone array noise suppression and speech enhancement methods typically operate in the time-frequency domain. Filter weights corresponding to the data received by each microphone array element are calculated within each time frame and frequency band, and then filtered to obtain the enhanced result. These commonly used microphone array noise suppression and speech enhancement methods currently utilize the correlation between microphone array channels to achieve excellent noise reduction performance with minimal speech distortion. They are a key technology for improving voice call quality and enhancing the accuracy of intelligent voice interaction.
[0003] The research results show that in the context of single-microphone noise reduction, exploiting the correlation between multi-frame data can improve noise reduction and speech distortion performance. Similar to the single-microphone case, in the context of microphone array noise suppression and speech enhancement, the simultaneous exploitation of the correlation between microphone channels and the correlation between multi-frame data can achieve better noise reduction performance than traditional methods that only utilize the correlation between microphone channels.
[0004] However, while performance improves, it also requires greater computing power. Conventional multi-frame expansion is generally performed on the original input data, which means that compared to single-frame data, multi-frame input data corresponds to an exponential or even multi-fold increase in computing power. How to minimize computing power consumption while performing multi-frame expansion is also a difficult problem for real-time processing systems. Summary of the Invention
[0005] The purpose of the present invention is to address the defects of the prior art and provide a method for generalized sidelobe elimination of a microphone array, so as to improve the generalized sidelobe elimination effect of the microphone array with the least possible computing power consumption.
[0006] The specific steps of the method of the present invention are:
[0007] Step (1) A microphone array including M microphones, the time domain received signal at the tth sampling point The superscript T indicates transpose. Represents the complex domain; where the received signal y of the mth microphone at the tth sampling point is m (t) = a m (t)*s(t)+n m (t), m=1,…,M, s(t) represents the speech signal, a m (t) represents the acoustic transfer function from the speech signal to the mth microphone, nm (t) represents the noise component received by the mth microphone, * represents the convolution operation;
[0008] Microphone array received signal in the short-time Fourier transform domain Where Y(k,l), X(k,l), and N(k,l) are M-dimensional vectors, which respectively represent the received signal spectrum, speech signal spectrum, and noise signal spectrum of the microphone array at the lth frequency point in the kth frame. The frequency domain transfer function A(k,l)=[A1(k,l) A2(k,l) … A M (k,l)] T , S(k,l) is the speech signal spectrum;
[0009] According to the frequency domain transfer function A(t,l), the fixed beamforming filter coefficient F(k,l) is obtained, and then the fixed beamforming filter output Y is obtained. FBF (k,l); fixed beamforming filter output Y FBF (k,l)=F H (k, l)Y(k, l); where F(k, l) is the fixed beamforming filter coefficient, the superscript H represents the conjugate transpose, and the filter coefficient is expressed as F(k, l) = A H (k,l).
[0010] Step (2) construct a sparse blocking matrix B(k,l) and obtain the blocking matrix output U(k,l) of the current frame;
[0011] Step (3) Expand the blocking matrix output U(k,l) of the current frame into a p-frame vector to obtain the p-frame reference noise signal U p (k, l), p = 2 to 8;
[0012] Step (4) is based on the p-frame reference noise signal U p (k,l) and the fixed beamforming filter output Y FBF (k,l), update the covariance matrix of the multi-frame reference noise signal And update the covariance matrix of the multi-frame reference noise signal and the fixed beamforming filter output
[0013] Step (5) is updated by and Obtain the coefficients H(k,l) of the adaptive interference cancellation filter;
[0014] Step (6) outputs the result Y of the sidelobe elimination filter MF (k,l).
[0015] Furthermore, step (2) is specifically:
[0016] Constructing a sparse blocking matrix The superscript * indicates the conjugation operation.
[0017] The blocking matrix output is the reference noise signal,
[0018] Furthermore, step (3) is specifically:
[0019] Expand the blocking matrix output U(k,l) of the current frame to a p-frame vector, then the p-frame reference noise signal Where U(k-p+1,l), U(k-1,l), and U(k,l) represent the reference noise signals corresponding to the lth frequency point of the k-p+1th frame, the k-1th frame, and the kth frame, respectively. If k < p, U(k-p+1,l) is a 0 vector.
[0020] Furthermore, step (4) is specifically:
[0021] initialization and is the unit matrix; update as follows:
[0022]
[0023]
[0024] The updated forgetting factor is 0<α<1.
[0025] Furthermore, step (5) is specifically as follows: the process of solving the adaptive interference cancellation filter can be regarded as the process of solving the Wiener filter solution, so the adaptive interference cancellation filter coefficient H(k,l) is finally obtained as The superscript -1 indicates a matrix inversion operation.
[0026] Further, in step (6), in, is the adaptive interference cancellation filter output,
[0027] The beneficial effects of the present invention are: in response to the defects of the prior art, a multi-frame expansion method for the generalized sidelobe elimination algorithm of the microphone array is proposed, which improves the effect of the generalized sidelobe elimination algorithm of the microphone array with the least possible computing power consumption. On the basis of the traditional microphone array generalized sidelobe elimination algorithm, the present invention constructs a sparse blocking matrix, utilizes the characteristic of reduced data dimension after the sparse blocking matrix, and outputs data in the blocking matrix, that is, the reference noise signal, to perform multi-frame expansion, thereby avoiding the problem of large amount of computing power consumption caused by multi-frame expansion on the most original input data, and also utilizes the correlation between multi-frame data to improve the performance of the microphone array generalized sidelobe elimination algorithm.
[0028] The advantages of this method are:
[0029] (1) Compared with the “single-frame input microphone array generalized sidelobe cancellation algorithm” method, the correlation between multi-frame data is utilized to improve the performance of the microphone array generalized sidelobe cancellation algorithm.
[0030] (2) Compared with the method of “generalized sidelobe elimination algorithm of microphone array with multi-frame expansion of the most original input data”, it does not need to calculate the multi-frame fixed beamforming and blocking filtering process, and because a sparse blocking matrix is used, the dimension of the blocking matrix output is smaller than the microphone array dimension, the overall calculation amount is small, and a lot of computing power consumption is saved. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic flow diagram of the present invention;
[0032] Figure 2 This is a block diagram of the generalized sidelobe elimination filter structure of the microphone array in the present invention. DETAILED DESCRIPTION
[0033] To facilitate understanding of the present invention, and to make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth in order to fully understand the present invention, and preferred embodiments of the present invention are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of the present invention. The present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited to the specific embodiments disclosed below.
[0034] A generalized sidelobe elimination method for microphone array, specifically Figure 1 shown.
[0035] Signal model: A microphone array consisting of M microphones, the time domain received signal at the tth sampling point The superscript T indicates transpose. Represents the complex domain; where the received signal y of the mth microphone at the tth sampling point is m (t) = a m (t)*s(t)+n m (t), m=1,…,M, s(t) represents the speech signal, a m (t) represents the acoustic transfer function from the speech signal to the mth microphone, n m(t) represents the noise component received by the m-th microphone, and * represents a convolution operation.
[0036] Microphone array received signal in the short-time Fourier transform domain Where Y(k,l), X(k,l), and N(k,l) are M-dimensional vectors, which respectively represent the received signal spectrum, speech signal spectrum, and noise signal spectrum of the microphone array at the lth frequency point in the kth frame. The frequency domain transfer function A(k,l)=[A1(k,l) A2(k,l) … A M (k,l)] T , S(k,l) is the spectrum of the speech signal.
[0037] In order to improve the robustness of the beamforming filter, the filter structure is split into two independent filters in the orthogonal subspace. The filter coefficient of the generalized sidelobe elimination method W(k,l)=F(k,l)-B(k,l)H(k,l) contains three parts, such as Figure 2 As shown: the first is the fixed beamforming filter F(k,l), the second is the blocking matrix B(k,l), and the third is the adaptive interference cancellation filter H(k,l).
[0038] Fixed beamforming filter: Fixed beamforming filter output Y FBF (k,l)=F H (k,l)Y(k,l); where F(k,l) is the fixed beamforming filter coefficient and the superscript H indicates the conjugate transpose.
[0039] The filter coefficient is expressed as F(k,l)=A H (k, l); If the fixed beamforming filter uses a delay-and-add filter structure, the frequency domain transfer function is expressed as The angular frequency ω corresponding to the lth frequency point l =2πf l , f l is the signal frequency corresponding to the lth frequency point, τ m is the time delay of the signal reaching the mth microphone, and j represents an imaginary number.
[0040] Taking a uniform linear array as an example, the frequency domain transfer function A(k,l) can be further expressed as:
[0041] A(k,l)=[1 e -jφ … e -j(M-1)φ ] T ,φ=2πf l d sinθ / v; where d is the uniform linear array spacing, θ is the incident angle of the speech signal relative to the microphone array, and f l is the signal frequency corresponding to the lth frequency point, and v is the voice propagation speed.
[0042] The purpose of the blocking matrix is to filter out the signal propagating along the acoustic path corresponding to the frequency transfer function A(k,l), so the column vector of the blocking matrix B(k,l) forms the null space of the frequency transfer function A(k,l), that is, it satisfies: B H (k,l)A(k,l)=0.
[0043] In order to maintain the filtering performance while reducing the computational complexity, a sparse blocking matrix is used, which is expressed as:
[0044] The superscript * indicates conjugation. The output of the blocking matrix is the reference noise signal. Since most of the speech components in the signal received by the microphone array are eliminated by the blocking matrix, most of the remaining components are noise signals, which can be expressed as:
[0045] In the traditional generalized sidelobe cancellation algorithm, the adaptive interference cancellation filter The input is the reference noise signal corresponding to the current frame k, and the corresponding output (where the superscript SF stands for single frame)
[0046] Because of the existence of inter-frame correlation, if multi-frame signal input is used, the inter-frame information of the signal can be reasonably utilized to further improve the filter effect. Therefore, we expand the input of the adaptive interference canceller to p-frame reference noise signals, p = 2 ~ 8, p-frame reference noise signals Among them, U(k,l), U(k-1,l),…, U(k-p+1,l) represent the noise reference signals corresponding to the kth frame, k-1th frame, and k-p+1th frame of the lth frequency point, that is, the output signal of the blocking matrix. At this time, the adaptive interference cancellation filter Output (where the superscript MF stands for Multi Frame)
[0047]
[0048] Output of the generalized sidelobe cancellation filter for multi-frame input
[0049] Energy spectral density of the filter output
[0050] in, is the energy spectral density of the fixed filter output, is the covariance matrix of the fixed filter output and the multi-frame output of the blocking matrix, Is the covariance matrix of the multi-frame output of the blocking matrix. Initialization and For the unit matrix, perform the following update:
[0051]
[0052] The updated forgetting factor is 0<α<1.
[0053] The process of solving the adaptive interference cancellation filter can be regarded as the process of solving the Wiener filter solution, so the adaptive interference cancellation filter coefficients are finally obtained. The superscript -1 indicates a matrix inversion operation.
[0054] The output result Y of the generalized sidelobe elimination filter with multi-frame input MF (k,l) is the final output.
Claims
1. A method for generalized sidelobe cancellation of a microphone array, characterized by: Step (1) A microphone array including M microphones, the time domain received signal at the tth sampling point The superscript T indicates transpose. Represents the complex domain; where the received signal y of the mth microphone at the tth sampling point is m (t) = a m (t)*s(t)+n m (t), m=1,…,M, s(t) represents the speech signal, a m (t) represents the acoustic transfer function from the speech signal to the mth microphone, n m (t) represents the noise component received by the mth microphone, * represents the convolution operation; Microphone array received signal in the short-time Fourier transform domain Among them, Y(k,l), X(k,l), and N(k,l) are M-dimensional vectors, which correspond to the received signal spectrum, speech signal spectrum, and noise signal spectrum of the microphone array at the lth frequency point in the kth frame, respectively. The frequency domain transfer function A(k,l)=[A1(k,l)A2(k,l)…A M (k,l)] T , S(k,l) is the one-dimensional channel speech signal spectrum; According to the frequency domain transfer function A(k,l), the fixed beamforming filter coefficient F(k,l) is obtained, and then the fixed beamforming filter output Y is obtained. FBF (k,l); fixed beamforming filter output Y FBF (k,l)=F H (k, l)Y(k, l); where F(k, l) is the fixed beamforming filter coefficient, the superscript H represents the conjugate transpose, and the filter coefficient is expressed as F(k, l) = A H (k,l); Step (2) construct a sparse blocking matrix B(k,l) and obtain the blocking matrix output U(k,l) of the current frame; Step (3) Expand the blocking matrix output U(k,l) of the current frame into a p-frame vector to obtain the p-frame reference noise signal U p (k, l), p = 2 to 8; Step (4) is based on the p-frame reference noise signal U p (k,l) and the fixed beamforming filter output Y FBF (k,l), update the covariance matrix of the multi-frame reference noise signal And update the covariance matrix of the multi-frame reference noise signal and the fixed beamforming filter output Step (5) is updated by and Obtain the coefficients H(k,l) of the adaptive interference cancellation filter; Step (6) outputs the result Y of the sidelobe elimination filter MF (k,l).
2. A microphone array generalized sidelobe elimination method according to claim 1, characterized in that: Step (2) is specifically: Constructing a sparse blocking matrix The superscript * indicates the conjugation operation; The blocking matrix output is the reference noise signal, 3. A microphone array generalized sidelobe elimination method as claimed in claim 2, characterized in that: Step (3) is specifically: Expand the blocking matrix output U(k,l) of the current frame to a p-frame vector, then the p-frame reference noise signal Where U(k-p+1,l), U(k-1,l), and U(k,l) represent the reference noise signals corresponding to the lth frequency point of the k-p+1th frame, the k-1th frame, and the kth frame, respectively. If k < p, U(k-p+1,l) is a 0 vector.
4. A microphone array generalized sidelobe elimination method as claimed in claim 3, characterized in that: Step (4) is specifically: initialization and is the unit matrix; update as follows: The updated forgetting factor is 0<α<1.
5. A microphone array generalized sidelobe elimination method as claimed in claim 4, characterized in that: In step (5), The superscript -1 indicates a matrix inversion operation.
6. A microphone array generalized sidelobe elimination method according to claim 5, characterized in that: In step (6), Adaptive interference cancellation filter output
Citation Information
Patent Citations
Voice signal enhancement system and method
CN102938254A
Robust GSC method based on coherence and energy ratio
CN111341340A