A low complexity cascaded generalized sidelobe canceling beamforming method and apparatus
By decomposing the high-dimensional generalized sidelobe canceller into a cascaded structure and combining it with multi-channel Wiener filtering and inter-frame iterative processing, the problems of high computational complexity and poor robustness in large microphone arrays are solved, achieving low-complexity interference noise suppression and target speech enhancement.
Patent Information
- Application Number
- CN202510023976.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Large microphone arrays suffer from high computational load, poor robustness, and inability to process data in real time when using existing beamforming techniques. In particular, the performance of adaptive beamforming technology deteriorates significantly when the number of array elements is large.
A cascaded generalized sidelobe cancellation beamforming method is adopted, which decomposes the high-dimensional generalized sidelobe canceller into two cascaded low-dimensional generalized sidelobe cancellation units. Combined with multi-channel Wiener filtering and inter-frame iterative processing, the beam weights are optimized through interleaving and fusion strategies, which reduces computational complexity and improves robustness.
It significantly reduces computational complexity, improves the real-time performance and robustness of the algorithm, effectively suppresses interference noise and enhances target speech, and is suitable for real-time applications with large microphone arrays.
Smart Images

Figure CN119851679B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of microphone array signal processing, in particular to a low-complexity cascaded generalized sidelobe canceller beam forming method and device. BACKGROUND
[0002] For a large microphone array with a large number of array elements, the commonly used beam forming techniques include fixed beam forming technique and adaptive beam forming technique. The fixed beam forming technique is simple to implement and has high robustness, but has the disadvantage of insufficient interference and noise suppression capability. The commonly used adaptive beam forming techniques include minimum variance distortion response (MVDR) beam forming, linearly constrained minimum variance (LCMV) beam forming, generalized sidelobe canceller (GSC) beam forming, etc., all of which have strong interference suppression capability, but have the problems of large amount of computation and poor robustness when the number of array elements is large, and especially the mismatch between array elements and the estimation deviation of noise covariance matrix may even cause performance deterioration. People have also proposed Kronecker product beam forming technique to reduce the dimension of the covariance matrix and improve the robustness, but at the cost of sacrificing the degrees of freedom, which may cause performance degradation. In practical applications, most of the schemes need to use noise estimation algorithm additionally, and the estimation accuracy is poor at low signal-to-noise ratio, which also reduces the robustness. In addition, the existing Kronecker beam forming does not involve a scheme for adapting to any array pattern based on frame processing. SUMMARY
[0003] In view of the deficiencies of the prior art, the present application aims to provide a low-complexity cascaded generalized sidelobe canceller beam forming method and device.
[0004] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme:
[0005] A low-complexity cascaded generalized sidelobe canceller beam forming method, comprising the following steps:
[0006] Step 1, using the coordinate information of each microphone element in the microphone array and the sound source wave direction parameter to obtain a constraint matrix, and using the microphone array to collect received time domain signals to obtain array frequency domain frame signals through frame division, windowing and short-time Fourier transform processing;
[0007] Step 2, interleaving the array frequency domain frame signals and the constraint matrix obtained in step 1 in multiple parallel branches to obtain the reordered array frequency domain frame signals and the constraint matrix of each branch respectively;
[0008] Step 3, setting multiple cascaded generalized sidelobe cancellers, each of which includes two low-dimensional front-stage generalized sidelobe cancellers and back-stage generalized sidelobe cancellers cascaded in front and back, updating the parameters through frame-based iterative operation to obtain the output frequency domain information of each branch;
[0009] Step 4: Combine the beam outputs of the cascaded generalized sidelobe cancellers of each branch to obtain the optimal solution of the beam output of the multi-channel cascaded generalized sidelobe cancellers;
[0010] The frequency domain output signals from steps 5 and 4 are then subjected to inverse short-time Fourier transform, windowing, and framing to obtain the time domain signal after noise interference suppression and target speech enhancement.
[0011] Furthermore, the specific process of step 1 is as follows:
[0012] Let the microphone array consist of M microphone elements, numbered sequentially as 1, 2, ..., m, ..., M. The spatial coordinates of the m-th microphone element are represented by r in a three-dimensional Cartesian coordinate system. m =[r x,m ,r y,m ,r z,m ] T T represents transpose, and the far-field incident angle azimuth parameters of the target sound source are known as follows: Where θ d and These are the elevation and azimuth angles of the target sound source d; assuming there are J interfering sound sources, the far-field incident angle and azimuth parameters of the j-th interfering sound source are... θ j and Let be the elevation angle and azimuth angle of the j-th interfering sound source, respectively, and then the constraint matrix is obtained:
[0013] C(k)=[v(k,Ω d ),v(k,Ω1),…,v(k,Ω J )] T k = 0, 1, 2, ... N fft / twenty one)
[0014] k is N fft Frequency index of the point Fourier transform, v(k,Ω) d ) and v(k,Ω j (j=1,2,…J) are the steering vectors of the target sound source and the interfering sound source, respectively, obtained using the microphone array element coordinates and the direction of the incoming sound wave:
[0015]
[0016] Where i represents the imaginary unit, κ d and κ j Here, are the wavenumbers of the target sound source and the interfering sound source, respectively; c is the speed of sound in space; and f is the wavenumber of the interfering sound source. k It is the frequency value corresponding to frequency point k;
[0017] Based on the distortion-free signal criterion, the array beam weight vector w is required to satisfy C.H w = g, g = [1, 0,... 0] T is a J+1 dimensional constraint target vector, C represents a constraint matrix, and H represents a conjugate transpose;
[0018] The array frequency domain frame signal of the lth frame and the kth frequency point obtained by framing, windowing, and short-time Fourier transform on the received time domain signal of the microphone array is x(k, l) = [x1(k, l),..., xM(k, l)]T. m (k, l),..., x M (k, l)] T .
[0019] Further, the specific process of step 2 is as follows:
[0020] P parallel interleaving processors are used for signal processing, and the interleaving processor interleaves and reorders the array frequency domain frame signal x(k, l) and the constraint matrix C(k). The interleaving table corresponding to the interleaving processor of the pth branch is m = I p (m'), and the array frequency domain frame signal output by the interleaving processor of the pth branch is:
[0021]
[0022] The array frequency domain frame signal vector x p (k, l) output by interleaving is obtained from the mth element of the input array frequency domain frame signal vector x(k, l); and the reordered interleaving output constraint matrix is:
[0023]
[0024] is the spatial position coordinate of the microphone element m;
[0025] The interleaving operation reorders the rows of the constraint matrix, that is, the steering vectors of the target sound source and the interference sound source are reordered using the corresponding interleaving table.
[0026] Further, the specific process of step 3 is as follows:
[0027] 3.1) The array frequency domain frame signal x p (k, l) and the constraint matrix C p (k) output by the interleaving of the pth branch are sent to the corresponding branch of the cascaded generalized sidelobe canceller, the input of the front-stage generalized sidelobe canceller is M1 dimensional, the input of the rear-stage generalized sidelobe canceller is M2 dimensional, and M = M1 x M2;
[0028] 3.2) Initialize the equivalent weight vector of the post-stage generalized sidelobe canceller w2(k,0) = w2(k,0), which is a M2-dimensional complex coefficient vector, l = 0 represents the 0th frame, i.e. the initial value, and the weight of the delay-sum beamformer is used as the initial value:
[0029]
[0030] 3.3) Calculate the pre-stage pretreatment matrix using w2(k,0) wherein the symbol represents the Kronecker product operation, is a M1-dimensional unit matrix, so Q1(k,0) is a MxM1-dimensional pretreatment matrix. The array frequency domain frame signal x p (k,0) output by the interleaving in step 2 is projected and transformed by the pre-stage pretreatment matrix to obtain a M1-dimensional frequency domain signal vector:
[0031]
[0032] H represents the conjugate transpose;
[0033] Meanwhile, the pretreatment is performed on the constraint matrix output by the interleaving to obtain the constraint matrix of the pre-stage generalized sidelobe canceller:
[0034]
[0035] 3.4) The pre-stage generalized sidelobe canceller performs weight optimization and signal enhancement on the pretreated M1-dimensional frequency domain signal vector;
[0036] The generalized sidelobe canceller is composed of a fixed beamformer on the main path, a blocking matrix and an adaptive noise canceller on the auxiliary path. The fixed beamformer on the main path allows the target signal to pass through and is enhanced to a certain extent. The fixed beam weight vector is located in the constraint subspace, and the fixed beam processing is performed on the input signal vector on the main path:
[0037]
[0038] wherein Y f,p1 (k,0) is the fixed beam output on the main path of the pre-stage generalized sidelobe canceller on the branch p, and ||·||2 represents the L2-norm of the vector;
[0039] The pre-stage blocking matrix B1 is located in the minimum variance subspace, and prevents the target signal from entering the auxiliary path. When the M1-dimensional frequency domain signal vector is sent into the auxiliary path, the vector u p1 (k,0) containing noise and interference information is obtained through the blocking matrix:
[0040]
[0041] The fixed beam output Y f,p1 (k, l) and the blocking matrix output u p1 (k, l) form an M1+1 dimensional input vector into the adaptive noise canceller, the weight vector of the adaptive noise canceller is updated frame by frame, the fixed beam output Y f,p1 (k, l) on the main path and the adaptive noise canceller output Y e,p1 (k, l) on the auxiliary path is:
[0042]
[0043] The weight coefficient is updated frame by frame using the multi-channel Wiener filtering method The iterative calculation formula can be expressed as:
[0044]
[0045] wherein μ is a regulation coefficient, Φ S,p1 (k, l), Φ N,p1 (k, l) and Φ SN,p1 (k, l) are the signal covariance matrix, the noise covariance matrix and the signal-noise mutual covariance vector of the front-stage generalized sidelobe canceller respectively, which are updated in each frame:
[0046]
[0047] wherein α is a smoothing coefficient, M1+1 dimensional signal vector h S,p1 (k, l) is composed of Y f,p1 (k, l) and M1 small constants ε close to 0, expressed as h S,p1 (k, l) = [Y f,p1 , ε, … ε] T ; M1+1 dimensional noise vector h N,p1 (k, l) is composed of the output vector and the input vector of the adaptive noise canceller, expressed as The superscript * represents the conjugate of the complex number; after each frame iteration, the front-stage equivalent weight vector is calculated for any frequency point:
[0048]
[0049] is the first element in the vector , h represents the vector composed of the second to last elements in the vector , the front-stage equivalent weight vector is sent into the rear-stage generalized sidelobe canceller;
[0050] 3.5) The post-stage generalized sidelobe canceller uses w1(k, l) to generate the post-stage preconditioning matrix where is an M2-dimensional identity matrix, so that Q2(k, l) is an M x M2-dimensional preconditioning matrix, and the array frequency domain frame signal x p (k, l) is projected by the post-stage preconditioning matrix to obtain an M2-dimensional frequency domain signal vector:
[0051]
[0052] Meanwhile, the post-stage preconditioning is performed on the constraint matrix of the interleaved output to obtain the constraint matrix of the post-stage generalized sidelobe canceller:
[0053]
[0054] 3.6) The fixed beamformer on the main path of the post-stage generalized sidelobe canceller performs fixed beam processing on the input signal vector:
[0055]
[0056] where Y f,p2 (k, l) is the fixed beam output on the main path of the post-stage generalized sidelobe canceller on the branch p;
[0057] The M2-dimensional frequency domain signal vector is processed by the post-stage blocking matrix B2 on the auxiliary path of the post-stage generalized sidelobe canceller to obtain a noise interference information vector u p2 (k, l) is expressed as:
[0058]
[0059] The fixed beam output Y f,p2 (k, l) and the blocking matrix output u p2 (k, l) are combined into an M2+1-dimensional input vector and are sent to the adaptive noise canceller, and the weight vector of the adaptive noise canceller is The multi-channel Wiener filtering method is used to update the weight coefficients frame by frame:
[0060]
[0061] Φ S,p2 (k, l), Φ N,p2 (k, l), and Φ SN,p2 (k, l) are the signal covariance matrix, the noise covariance matrix, and the signal-noise cross-covariance vector of the post-stage generalized sidelobe canceller, respectively, and the three are updated in each frame:
[0062]
[0063] where α = 0.02, h S,p2 (k, l) = [Y f,p2 , ε,... ε] T , For any frequency point, the equivalent weight vector of the latter stage is calculated after each frame iteration:
[0064]
[0065] where, is the first element in the vector , and represents the vector composed of the second to last elements in the vector .
[0066] 3.7) The beam output of the cascaded generalized sidelobe canceller can be represented as Y f,p2 (k, l) and the difference between Y e,p2 (k, l) and Y
[0067]
[0068] When all frequency points k are traversed, the total beam output on branch p is obtained.
[0069] Further, the specific process of step 4 is as follows:
[0070] The beam outputs Y p (k, l), p = 1,..., P of each branch are fused using the linearly constrained minimum variance criterion. Since the target signal distortionless condition is satisfied, the branch with the minimum output power in the target direction is considered to be the optimal one. The branch number corresponding to the optimal solution obtained by fusing each frame signal is :
[0071]
[0072] In the formula, the objective function corresponding to the frequency point on each frame signal is:
[0073] Γ p (k, l) = (1 - β)Γ p (k, l - 1) + β | Y p (k, l) | 2 (24)
[0074] where β is the smoothing coefficient, and β = 0.015 is taken, then is determined as the optimal solution of the beam forming output on the kth frequency point of the lth frame.
[0075] Further, in steps 3.4) and 3.6, μ = 1.2 and α = 0.02.
[0076] The application also provides a low-complexity cascaded generalized sidelobe canceling beam forming device for implementing the method, comprising a microphone array, a microphone array signal processor, a signal fusion processor, a signal enhancement processor, a plurality of interleaving processors and a plurality of cascaded generalized sidelobe canceling processors, each of the cascaded generalized sidelobe canceling processors comprises a front-stage generalized sidelobe canceling processor and a rear-stage generalized sidelobe canceling processor connected in cascade; the front-stage generalized sidelobe canceling processor of each of the cascaded generalized sidelobe canceling processors is communicatively connected with one of the interleaving processors, each of the interleaving processors is communicatively connected with the microphone array signal processor, and the microphone array is communicatively connected with the microphone array signal processor; the rear-stage generalized sidelobe canceling processor of each of the cascaded generalized sidelobe canceling processors is communicatively connected with the signal fusion processor, and the signal fusion processor is communicatively connected with the signal enhancement processor.
[0077] The microphone array signal processor is used for processing the time-domain signals received by the microphone array to obtain array frequency-domain frame signals.
[0078] The interleaving processor is used for interleaving the array frequency-domain frame signals.
[0079] The front-stage generalized sidelobe canceling processor and the rear-stage generalized sidelobe canceling processor are both used for performing generalized sidelobe canceling processing on the array frequency-domain frame signals output by the interleaving processor, and the front-stage generalized sidelobe canceling processor and the rear-stage generalized sidelobe canceling processor transmit equivalent weight vector information to each other for iterative updating.
[0080] The signal fusion processor is used for fusing the beams output by all the rear-stage generalized sidelobe canceling processors.
[0081] The signal enhancement processor is used for enhancing the signals output by the signal fusion processor.
[0082] Furthermore, the front-stage generalized sidelobe canceling processor and the rear-stage generalized sidelobe canceling processor both comprise a main path and an auxiliary path, the main path comprises a fixed beam former, the auxiliary path comprises a blocking matrix and an adaptive noise canceling device, and the blocking matrix is composed of interference filters.
[0083] The application has the following advantages:
[0084] 1. The beam former with the cascaded generalized sidelobe canceling structure can decompose a high-dimension generalized sidelobe canceling device into two low-dimension generalized sidelobe canceling units connected in cascade, thereby significantly reducing the operation complexity and improving the robustness, and the frame-based front-stage and rear-stage feedback iterative processing mode can improve the real-time performance of the algorithm and has higher practical value.
[0085] 2, The application utilizes the idea of multi-channel Wiener filtering to construct an interference noise cancellation module, directly updates the filter coefficients between frames, without additional noise estimation, simple implementation and high robustness, and can balance the interference noise suppression and target speech distortion.
[0086] 3, The application can improve the degree of freedom of beam weight coefficient search through multi-branch fusion and interleaving strategy, which is more conducive to the algorithm to find the optimal solution, and improves the performance of target enhancement and interference noise suppression.
[0087] In summary, the application solves the problems of poor performance, low robustness and large computation of existing adaptive beamforming technology methods for large microphone arrays by using the cascade generalized sidelobe cancellation structure based on frame iteration processing, multi-channel Wiener filtering interference noise cancellation and multi-branch fusion and interleaving strategy. BRIEF DESCRIPTION OF DRAWINGS
[0088] Figure 1 The overall flowchart of the method in Example 1 of the application is shown in the figure.
[0089] Figure 2 The array topology and signal model diagram in the method of Example 1 of the application is shown in the figure.
[0090] Figure 3 The flowchart of the cascade generalized sidelobe cancellation method in the method of Example 1 of the application is shown in the figure.
[0091] Figure 4 The signal output power pattern of each beamforming method in the experiment of Example 2 of the application is shown in the figure.
[0092] Figure 5 The output gain comparison chart of each beamforming method under different signal interference ratios in the experiment of Example 2 of the application is shown in the figure.
[0093] Figure 6 The operation time simulation comparison of each beamforming method in the experiment of Example 2 of the application is shown in the figure.
[0094] Figure 7 The structure schematic diagram of the device in Example 3 of the application is shown in the figure. DETAILED DESCRIPTION
[0095] The application will be further described below with reference to the accompanying drawings, and it should be noted that the embodiments are based on the technical solution, and detailed implementation and specific operation processes are given, but the protection scope of the application is not limited to the embodiments.
[0096] Example 1
[0097] The embodiment provides a low-complexity cascade generalized sidelobe cancellation beamforming method, as shown in the figure, including the following steps: Figure 1
[0098] Step 1, using the coordinate information of each microphone element in the microphone array and the sound source wave direction parameter to obtain the constraint matrix, the time domain signal received by the microphone array is processed by frame division, windowing and short time Fourier transform to obtain the array frequency domain frame signal.
[0099] The method of the embodiment is applicable to any array type of microphone array, including linear array, surface array and spherical array, etc. The embodiment is described with a three-dimensional microphone array with microphone elements distributed in space, and the array topology and signal model are shown in Figure 2 The microphone array is composed of M omnidirectional microphone elements, numbered 1, 2, … m, …, M, and the spatial position coordinates of the mth omnidirectional microphone element are represented as r m = [r x,m , r y,m , r z,m ] T in the three-dimensional Cartesian coordinate system, and the far-field incident angle azimuth parameter of the target sound source is known as where θ d and are the pitch angle and azimuth angle of the target sound source d. There are J interference sound sources, and the far-field incident angle azimuth parameter of the jth interference sound source is θ j and are the pitch angle and azimuth angle of the jth interference sound source, and the constraint matrix is obtained:
[0100] C(k) = [v(k, Ω d ), v(k, Ω1), …, v(k, Ω J )] T k = 0, 1, 2, … N fft / 2 (1)
[0101] k is the frequency point number of N fft point Fourier transform, v(k, Ω d ) and v(k, Ω j )(j = 1, 2, … J) are the steering vectors of the target sound source and the interference sound source, which are obtained using the microphone element coordinates and the sound source wave direction information:
[0102]
[0103] where i represents the imaginary unit, κ d and κ j are the sound source wave numbers of the target sound source and the interference sound source, c is the propagation speed of sound in space, and f k is the frequency value corresponding to the frequency point k.
[0104] Based on the signal distortionless criterion, the array beam weight vector w generally requires to satisfy C H w = g, g = [1, 0, … 0] T is a J+1 dimensional constraint target vector, C represents a constraint matrix, and H represents a conjugate transpose.
[0105] The array frequency domain frame signal of the lth frame and the kth frequency point obtained by framing, windowing and short-time Fourier transform on the received time domain signal of the microphone array is x(k, l) = [x1(k, l), …, xM(k, l)]T. m (k, l), …, x M (k, l)] T .
[0106] Step 2, the array frequency domain frame signal obtained in step 1 and the constraint matrix are interleaved in multiple parallel branches to obtain the reordered array frequency domain frame signal and the constraint matrix of each branch, respectively.
[0107] The method of the embodiment adopts P parallel interleaving processors for signal processing, thereby providing more degrees of freedom for subsequent beam weight optimization. The interleaving processor interleaves and reorders the array frequency domain frame signal x(k, l) and the constraint matrix C(k). The interleaving table corresponding to the interleaving processor of the pth branch is m = I p (m'), then the interleaving output array frequency domain frame signal of the pth branch is:
[0108]
[0109] The interleaving output array frequency domain frame signal vector x p (k, l) is taken from the mth element of the input array frequency domain frame signal vector x(k, l). The reordered interleaving output constraint matrix is:
[0110]
[0111] is the spatial position coordinate of the microphone element m;
[0112] The interleaving operation reorders the rows of the constraint matrix, that is, the steering vectors of the target sound source and the interference sound source are reordered by the corresponding interleaving table.
[0113] Step 3, the method of the embodiment decomposes the high-dimensional generalized sidelobe canceller into a cascaded generalized sidelobe canceller, which includes two low-dimensional front-stage generalized sidelobe cancellers and rear-stage generalized sidelobe cancellers cascaded in front and back, and obtains the output frequency domain information of each branch by updating the parameters based on frame iteration operation. As shown in the following formula (5), the specific process is as follows: Figure 3 .
[0114] 3.1) The array frequency domain frame signal x of the branch p interlaced output p (k, l) and the constraint matrix C p (k) is sent to the cascaded generalized sidelobe canceller of the corresponding branch, the input of the former generalized sidelobe canceller is M1 dimension, and the input of the latter generalized sidelobe canceller is M2 dimension, M = M1 x M2.
[0115] 3.2) The equivalent weight vector w2(k, l) = w2(k, 0) of the latter generalized sidelobe canceller is initialized, which is an M2 dimension complex coefficient vector, and l = 0 represents the 0th frame, that is, the initial value. In order to ensure the reliability of the algorithm, the weight of the delay-sum beamforming is used as the initial value:
[0116]
[0117] It can be seen that the initialization value is only related to the target sound source steering vector information of the first M2 microphone array elements after interlacing.
[0118] 3.3) The former pre-processing matrix is calculated and generated by using w2(k, l) Wherein the symbol represents the Kronecker product operation, is an M1 dimension unit matrix, so Q1(k, l) is an M x M1 dimension pre-processing matrix, and the array frequency domain frame signal x p (k, l) of the interlaced output in step 2 is projected and transformed by the former pre-processing matrix to obtain an M1 dimension frequency domain signal vector:
[0119]
[0120] H represents the conjugate transpose;
[0121] At the same time, the constraint matrix of the interlaced output is pre-processed to obtain the constraint matrix of the former generalized sidelobe canceller:
[0122]
[0123] 3.4) The former generalized sidelobe canceller optimizes the weight and enhances the signal of the pre-processed M1 dimension frequency domain signal vector.
[0124] The generalized sidelobe canceller is composed of a fixed beamformer on the main road, a blocking matrix on the auxiliary road, and an adaptive noise canceller. The fixed beamformer on the main road makes the target signal pass through and is enhanced to a certain extent. The fixed beam weight vector is located in the constraint subspace, and the fixed beam processing is performed on the input signal vector on the main road:
[0125]
[0126] where Y f,p1 (k, l) is the fixed beam output on the main path of the preceding generalized sidelobe canceller on the branch p, and || · ||2 denotes the L2-norm of a vector.
[0127] The preceding blocking matrix B1 is located in the minimum variance subspace, which prevents the target signal from entering the auxiliary path. The M1-dimensional frequency domain signal vector is sent into the auxiliary path through the blocking matrix to obtain a vector u p1 (k, l) containing noise and interference information.
[0128]
[0129] The fixed beam output Y f,p1 (k, l) and the blocking matrix output u p1 (k, l) form an M1+1-dimensional input vector which is sent into the adaptive noise canceller. The weight vector of the adaptive noise canceller is updated frame by frame. The difference between the fixed beam output Y f,p1 (k, l) on the main path and the adaptive noise canceller output Y e,p1 (k, l) on the auxiliary path is:
[0130]
[0131] The weight coefficient is updated frame by frame using the multi-channel Wiener filtering method. The iterative calculation formula can be expressed as:
[0132]
[0133] where μ is a regulation coefficient, and here μ = 1.2; Φ S,p1 (k, l), Φ N,p1 (k, l), and Φ SN,p1 (k, l) are the signal covariance matrix, the noise covariance matrix, and the signal-noise mutual covariance vector of the preceding generalized sidelobe canceller, respectively, which are updated in each frame:
[0134]
[0135] where α is a smoothing coefficient, and α = 0.02. The M1+1-dimensional signal vector h S,p1 (k, l) is composed of Y f,p1 (k, l) and M1 small constants ε (such as ε = 10 -10 ) close to 0, and is expressed as h S,p1 (k, l) = [Y f,p1 , ε, …, ε] T ; and the M1+1-dimensional noise vector h N,p1(k, l) consists of the output vector of the adaptive noise canceller and the input vector, denoted as The superscript * denotes taking the conjugate of a complex number. For any frequency point, the pre-stage equivalent weight vector is calculated after each frame iteration:
[0136]
[0137] The first element of the vector The second to last element of the vector
[0138] 3.5) The post-stage generalized sidelobe canceller calculates the post-stage pre-processing matrix where is an M2-dimensional identity matrix, so Q2(k, l) is an M x M2-dimensional pre-processing matrix. The array frequency domain frame signal x p (k, l) is projected by the post-stage pre-processing matrix to obtain an M2-dimensional frequency domain signal vector:
[0139]
[0140] At the same time, the post-stage pre-processing of the constraint matrix of the interleaved output is performed to obtain the constraint matrix of the post-stage generalized sidelobe canceller:
[0141]
[0142] 3.6) The fixed beamformer on the main path of the post-stage generalized sidelobe canceller performs fixed beam processing on the input signal vector:
[0143]
[0144] where Y f,p2 (k, l) is the fixed beam output on the main path of the post-stage generalized sidelobe canceller on the branch p;
[0145] The M2-dimensional frequency domain signal vector is processed by the post-stage blocking matrix B2 on the auxiliary path of the post-stage generalized sidelobe canceller to obtain a noise interference information vector u p2 (k, l) is denoted as:
[0146]
[0147] The M2+1-dimensional input vector f,p2 (k, l) is composed of the fixed beam output Y p2 (k, l) and the blocking matrix output u and input into the adaptive noise canceller, the weight vector of the adaptive noise canceller is The weight coefficient is updated frame by frame by using the multi-channel Wiener filtering method:
[0148]
[0149] Similarly, μ = 1.2, Φ S,p2 (k, l), Φ N,p2 (k, l) and Φ SN,p2 (k, l) are the signal covariance matrix, noise covariance matrix and signal-noise cross-covariance vector of the post-stage generalized sidelobe canceller, respectively, which are updated in each frame:
[0150]
[0151] where α = 0.02, h S,p2 (k, l) = [Y f,p2 , ε, … ε] T , After each frame iteration, the post-stage equivalent weight vector is calculated for any frequency point:
[0152]
[0153] where, is the first element in the vector , and the vector composed of the second to last elements in the vector .
[0154] 3.7) The beam output of the cascaded generalized sidelobe canceller can be represented as Y f,p2 (k, l) and Y e,p2 (k, l) are:
[0155]
[0156] When all frequency points k are traversed, the total beam output on the branch p is obtained.
[0157] Step 4, fuse the beam outputs of the cascaded generalized sidelobe cancellers of each branch to obtain the optimal solution of the beam output of the multi-path cascaded generalized sidelobe canceller.
[0158] The beam output Y pThe constraint matrix in each branch is consistent in essence although it is interleaved, and since the cascade generalized sidelobe canceller structure of each branch and the adjusting coefficient μ are the same, it can be considered that each branch is very close to achieving the linear constraint condition in the target, and the main reason for inconsistency is the difference in noise interference suppression amount, and since the target signal distortionless condition is met, it can be considered that the branch with the minimum output power of the target direction is optimal, and the optimal solution corresponding to the branch number is obtained by fusing each frame signal For:
[0159]
[0160] In the formula, the target function corresponding to each frame signal on the frequency point is:
[0161] Γ p (k, l) = (1 - β)Γ p (k, l - 1) + β | Y p (k, l)| 2 (24)
[0162] Wherein, β is a smoothing coefficient, and β = 0.015 is taken, and then it is determined that is the beamforming output optimal solution of the lth frame and the kth frequency point.
[0163] Step 5, the frequency domain output signal is subjected to inverse short-time Fourier transform, windowing and framing to obtain a time domain signal after noise interference suppression and target speech enhancement.
[0164] The frequency domain output signal The target sound time domain output signal can be restored through the inverse short-time Fourier transform, windowing and framing operations. When the direction of the target sound source or the interference sound source changes, it is necessary to return to step 1 to generate a new constraint matrix, initialize and restart the entire calculation process.
[0165] Embodiment 2
[0166] This embodiment aims to verify the effectiveness of the method of embodiment 1 through experiments.
[0167] In the experiment, the number of microphone array elements is M = 20, the cascade generalized sidelobe canceller takes M1 = 5 and M2 = 4, the parameter definition is adjusting coefficient μ = 1.2, smoothing factor α = 0.02, β = 0.015, and there is a target sound source and an interference sound source in the simulation environment.
[0168] Figure 4To be the embodiment 1 method and the comparison of the common beam forming method with the wideband signal output power pattern changing with the incident angle, it can be seen that the traditional fixed beam forming technology can obtain low sidelobe by windowing method while effectively enhancing the target signal, but it cannot eliminate the directional interference sound source, the adaptive beam forming method, the Kronecker product minimum variance distortionless response beam forming KPMVDR method and the generalized sidelobe cancellation GSC method can form null in the direction of the interference sound source, effectively suppress the signal from the interference direction, but at the same time, the sidelobe will be lifted, the directivity at low frequency will be worse, leading to the introduction of directional interference noise. It can be seen that the embodiment 1 method can suppress the interference sound source while the sidelobe is much lower than the KPMVDR and GSC methods, and is close to the sidelobe suppression fixed beam forming method, the out-of-main-lobe sidelobe suppression degree can reach about 20dB, and the multi-channel wiener filter interference noise elimination and multi-branch fusion strategy can significantly improve the interference and noise suppression degree and robustness, and improve the target speech quality.
[0169] Figure 5 It is the output gain graph comparison of the beam forming method under different signal interference ratios, there are three groups of interference sound sources in the environment, including white noise, indoor noisy noise and street noise, with the change of interference intensity, the wideband signal gain of the traditional delay sum fixed beam method (DS), MVDR method, KPMVDR method, GSC method and embodiment 1 method is compared, obviously, in the case of low signal interference ratio, the performance of delay sum beam former is poor, and the performance gradually improves with the increase of signal interference ratio, which shows that in the condition of high interference and noise, the fixed beam former has great limitation. At the same time, whether in low interference or high interference level, the performance of KPMVDR method, GSC method and embodiment 1 method beam former is much better than that of minimum variance distortionless response beam former, and the processing gain of embodiment 1 method is slightly better than that of KPMVDR method and GSC method due to the full use of array degree of freedom and the precise estimation of noise statistical characteristics by potential adaptive multi-channel wiener filter canceller, which verifies the robustness and effectiveness of embodiment 1 method under different interference and noise scenes.
[0170] Figure 6The computation time of the generalized sidelobe cancellation method and the cascaded generalized sidelobe cancellation method in Example 1 were compared. The simulation used an array with M array elements (M1 = M2 = 3, 4, ... 10) and simulated the time required to process a 30-second audio signal on MATLAB. It can be seen that the computation time of the GSC method increases significantly with the number of array elements. This is because as the number of array elements increases, high-dimensional matrix operations are required, resulting in increasing complexity, making it unsuitable for scenarios with high real-time requirements. In contrast, the method in Example 1 is not sensitive to the number of array elements, especially for large arrays with a large number of elements. The method in Example 1 decomposes the high-dimensional generalized sidelobe canceller into two low-dimensional generalized sidelobe cancellers through cascading, greatly reducing the computation time and complexity, making it more conducive to real-time implementation and application deployment.
[0171] Example 3
[0172] This embodiment provides a low-complexity cascaded generalized sidelobe phase-cancellation beamforming device, such as... Figure 7 As shown, the system includes a microphone array, a microphone array signal processor, a signal fusion processor, a signal enhancement processor, multiple interleaving processors, and multiple cascaded generalized sidelobe cancellation processors. Each cascaded generalized sidelobe cancellation processor includes a pre-stage generalized sidelobe cancellation processor and a post-stage generalized sidelobe cancellation processor cascaded together. The pre-stage generalized sidelobe cancellation processor of each cascaded generalized sidelobe cancellation processor is communicatively connected to an interleaving processor, and each interleaving processor is communicatively connected to the microphone array signal processor. The microphone array is communicatively connected to the microphone array signal processor. The post-stage generalized sidelobe cancellation processor of each cascaded generalized sidelobe cancellation processor is communicatively connected to the signal fusion processor, and the signal fusion processor is communicatively connected to the signal enhancement processor.
[0173] The microphone array signal processor is used to process the time-domain signals acquired and received by the microphone array to obtain array frequency-domain frame signals;
[0174] The interleaving processor is used to interleave array frequency domain frame signals;
[0175] Both the pre-stage generalized sidelobe cancellation processor and the post-stage generalized sidelobe cancellation processor are used to perform generalized sidelobe cancellation processing on the array frequency domain frame signal output by the interleaving processor; the pre-stage generalized sidelobe cancellation processor and the post-stage generalized sidelobe cancellation processor exchange equivalent weight vector information for iterative updates;
[0176] The signal fusion processor is used to fuse the beams output by all subsequent generalized sidelobe cancellation processors;
[0177] The signal enhancement processor is used to enhance the signal output by the signal fusion processor.
[0178] In the embodiment, the front-stage generalized sidelobe canceling processor and the rear-stage generalized sidelobe canceling processor each include a main path and an auxiliary path, the main path includes a fixed beamformer, and the auxiliary path includes a blocking matrix and an adaptive noise canceller; the blocking matrix is composed of interference filters.
[0179] For those skilled in the art, various corresponding changes and modifications can be given according to the above technical solutions and concepts, and all these changes and modifications should be included in the protection scope of the claims of the present application.
Claims
1. A low complexity cascaded generalized sidelobe canceling beamforming method, characterized in that, The method comprises the following steps: Step 1, using the coordinate information of each microphone element in the microphone array and the sound source direction parameter to obtain a constraint matrix, the time domain signal received by the microphone array is processed by frame division, windowing and short-time Fourier transform to obtain an array frequency domain frame signal; The microphone array is composed of M microphone elements, numbered in turn as 1, 2,...m,...,M, and the spatial position coordinates of the mth microphone element are represented in a three-dimensional Cartesian coordinate system as r m = [r x,m , r y,m , r z,m ] T , T represents transposition, and the far-field incident angle azimuth parameter of the target sound source is known as where θ d and are the pitch angle and azimuth angle of the target sound source d; there are J interference sound sources, and the far-field incident angle azimuth parameter of the jth interference sound source is j = 1, 2,...J, θ j and are the pitch angle and azimuth angle of the jth interference sound source, and the constraint matrix is obtained: C(k) = [v(k, Ω d ), v(k, Ω J ),..., v(k, Ω T )] fft k = 0, 1, 2,... N / 2 (1) k is N fft The frequency point sequence number of point Fourier transform, v(k, Ω d ) and v(k, Ω j (j = 1, 2, … J) are the steering vectors of the target sound source and the interference sound source respectively, which are obtained by using the microphone array element coordinates and the sound source wave direction information: where i represents the imaginary unit, k d and k j are the source wavenumbers of the target and interfering sources, respectively, c is the speed of sound in space, f k is the frequency value corresponding to the frequency point k; Based on the signal distortionless criterion, the array beam weight vector w is required to satisfy C H w = g, g = [1, 0,... 0] T is a J+1 dimensional constraint target vector, C represents a constraint matrix, and H represents a conjugate transpose. The array frequency domain signal of the lth frame and the kth frequency point obtained by framing, windowing and short-time Fourier transform of the received time domain signal collected by the microphone array is x(k, l) = [x1(k, l),..., xM(k, l)]T. m (k, l),..., x M (k, l)] T ; Step 2, the array frequency domain frame signal and the constraint matrix obtained in step 1 are interleaved in multiple parallel branches to obtain the reordered array frequency domain frame signal and the constraint matrix of each branch respectively; Step 3, multiple cascaded generalized sidelobe cancellers are set, each cascaded generalized sidelobe canceller comprises two low-dimensional front-stage and rear-stage generalized sidelobe cancellers which are cascaded in sequence, the parameters are updated through frame-based iterative operation to obtain the output frequency domain information of each branch; Step 4, the beam outputs of the cascaded generalized sidelobe cancellers of each branch are fused to obtain the optimal solution of the beam output of the multiple cascaded generalized sidelobe cancellers; Step 5, the frequency domain output signal finally output by step 4 is subjected to inverse short-time Fourier transform, windowing and frame grouping to obtain a time domain signal after noise interference suppression and target speech enhancement.
2. The method of claim 1, wherein, The specific process of step 2 is: P parallel interleaving processors are used for signal processing, the interleaving processors reorder the array frequency domain frame signal x(k, l) and the constraint matrix C(k), the interleaving table corresponding to the interleaving processor of the pth branch is m = I p (m'), then the array frequency domain frame signal interleaved and output by the interleaving processor of the pth branch is the array frequency domain frame signal vector x of the interleaved output p the m'th element of (k, l) is taken from the m'th element of the input array frequency domain frame signal vector x(k, l); the reordered interleaved output constraint matrix is is the spatial position coordinate for the microphone array element m; The interleaving operation reorders the rows of the constraint matrix, that is, the steering vectors of the target sound source and the interference sound source are reordered by using the corresponding interleaving table.
3. The method of claim 2, wherein, The specific process of step 3 is: 3.1) Array frequency domain frame signal x of branch p interleave output p (k, l) and constraint matrix C p (k) is sent into the corresponding branch of the cascade generalized sidelobe canceller, the input of the former generalized sidelobe canceller is M1 dimension, the input of the latter generalized sidelobe canceller is M2 dimension, M = M1 x M2; 3.2) The equivalent weight vector w2(k, l) of the rear-stage generalized sidelobe canceller is initialized as w2(k, 0), which is an M2-dimensional complex coefficient vector, and l=0 represents the initial value of the 0th frame, and the weight value of the delay-sum beam forming is used as the initial value: 3.3) Compute the pre-processing matrix using w2(k,l) where the symbol denotes the Kronecker product operation, is the M1-dimensional identity matrix, so that Q1(k,l) is an M x M1-dimensional pre-processing matrix, and the interleaved output array- frequency frame signal x p (k,l) is projected by the pre-processing matrix to obtain the M1-dimensional frequency signal vector H represents conjugate transpose; Meanwhile, the constraint matrix output by interleaving is preprocessed to obtain the constraint matrix of the front-stage generalized sidelobe canceller: 3.4) The front-stage generalized sidelobe canceller optimizes the weight value and enhances the signal of the M1-dimensional frequency domain signal vector obtained by preprocessing; The generalized sidelobe canceller consists of three parts: a fixed beamformer in the main path, a blocking matrix and an adaptive noise canceller in the auxiliary path. The fixed beamformer in the main path passes the target signal and gets a certain degree of enhancement. The fixed beam weight vector is located in the constraint subspace, and the fixed beam processing is performed on the input signal vector in the main path: where Y f,p1 (k, l) is the fixed beam output on the main path of the preceding generalized sidelobe canceller on branch p, and || · ||2denotes the L2-norm of a vector. The pre-stage blocking matrix B1 is located in the minimum variance subspace, which prevents the target signal from entering the auxiliary path. The M1-dimensional frequency domain signal vector is sent into the auxiliary path through the blocking matrix to obtain a vector u containing noise and interference information p1 (k, l): The fixed beam output Y f,p1 (k, l) and the blocking matrix output u p1 (k, l) form an M1+1 dimensional input vector to the adaptive noise canceller, whose weight vector is updated frame by frame, the difference between the fixed beam output Y f,p1 (k, l) on the main path and the adaptive noise canceller output Y e,p1 (k, l) on the secondary path is: The multi-channel wiener filtering method is used to update the weight coefficient frame by frame The iterative calculation formula can be expressed as: where μ is a scaling factor, Φ S,p1 (k, l), Φ N,p1 (k, l), and Φ SN,p1 (k, l) are the signal covariance matrix, the noise covariance matrix, and the signal-noise cross-covariance vector of the pre-stage generalized sidelobe canceller, respectively, which are updated in each frame: where a is a smoothing factor, M1+1 dimensional signal vector h S,p1 (k, l) is composed of Y f,p1 (k, l) and M1 small constants close to 0, denoted as h S,p1 (k, l) = [Y f,p1 , ε,... ε T ; M1+1 dimensional noise vector h N,p1 (k, l) is composed of the output vector and the input vector of the adaptive noise canceller, denoted as The superscript * denotes taking the conjugate of a complex number; for any frequency point, the equivalent weight vector of the previous stage is calculated after each frame iteration: For vectors The first element in Representing vectors The vector consisting of the second to the last element in the first stage, the equivalent weight vector of the previous stage, is fed into the generalized sidelobe canceller of the next stage. 3.5) The post-stage generalized sidelobe canceller uses w2(k, l) to compute the post-stage pre-processing matrix where is an M2-dimensional identity matrix, so Q2(k, l) is an M x M2-dimensional pre-processing matrix, and the interleaved output of the array frequency domain frame signal x p (k, l) is projected by the post-stage pre-processing matrix to obtain an M2-dimensional frequency domain signal vector: Meanwhile, the constraint matrix output by interleaving is preprocessed to obtain the constraint matrix of the rear-stage generalized sidelobe canceller: 3.6) The fixed beam former on the main path of the rear-stage generalized sidelobe canceller performs fixed beam processing on the input signal vector: where Y f,p2 (k, l) is the fixed beam output on the main path of the post-stage generalized sidelobe canceller on the branch p; The M2-dimensional frequency domain signal vector is processed by the post-stage blocking matrix B2 on the auxiliary path of the post-stage generalized sidelobe canceller to obtain a noise interference information vector u p2 (k, l) is represented as: The fixed beam output Y f,p2 (k, l) and the blocking matrix output u p2 (k, l) to form an M2+1 dimensional input vector and sent to the adaptive noise canceller, the weight vector of the adaptive noise canceller is The weight coefficient is updated frame by frame by using the multi-channel Wiener filtering method: Φ S,p2 (k, l), Φ N,p2 (k, l) and Φ SN,p2 (k, l) are the signal covariance matrix, the noise covariance matrix and the signal-noise cross-covariance vector of the post-stage generalized sidelobe canceller, respectively, which are updated in each frame: where a = 0.02, h S,p2 (k, l) = [Y f,p2 , ε,... ε T , For any frequency point, the equivalent weight vector of the latter stage is calculated after each frame iteration ends: wherein is a vector of the first element, denotes a vector of the second to last element. 3.7) The beam output of a cascaded generalized sidelobe canceller can be represented as Y f,p2 (k, l) and Y e,p2 (k, l) difference: After traversing all frequency points k, the total beam output on the branch p is obtained.
4. The method of claim 3, wherein, The specific process of step 4 is: The linear constrained minimum variance criterion is used to fuse the beam outputs Y p (k, l), p = 1, …, P, and the branch with the minimum output power in the target direction is considered to be optimal, and the branch sequence number corresponding to the optimal solution is obtained by fusing each frame of signals is: In the formula, the target function on the corresponding frequency point of each frame signal is: Γ p (k, l) = (1 - β)Γ p (k, l - 1) + β|Y p (k, l)| 2 (24) Wherein, β is a smoothing coefficient, take β = 0.015, then determine is the optimal solution of the beamforming output on the lth frame and the kth frequency point.
5. The method of claim 3, wherein, In steps 3.4) and 3.6, μ=1.2 and α=0.
02.
6. A low complexity cascaded generalized sidelobe canceller beamforming apparatus for implementing the method of any one of claims 1-5, characterized by The microphone array, the microphone array signal processor, the signal fusion processor, the signal enhancement processor, multiple interleaving processors and multiple cascaded generalized sidelobe cancellation processors are included, each cascaded generalized sidelobe cancellation processor comprises a front-stage generalized sidelobe cancellation processor and a rear-stage generalized sidelobe cancellation processor which are cascaded in sequence; the front-stage generalized sidelobe cancellation processor of each cascaded generalized sidelobe cancellation processor is in communication connection with an interleaving processor, each interleaving processor is in communication connection with the microphone array signal processor, the microphone array is in communication connection with the microphone array signal processor; the rear-stage generalized sidelobe cancellation processor of each cascaded generalized sidelobe cancellation processor is in communication connection with the signal fusion processor, and the signal fusion processor is in communication connection with the signal enhancement processor; The microphone array signal processor is configured to process a time domain signal received by a microphone array to obtain an array frequency domain frame signal; The interleaving processor is configured to perform interleaving processing on the array frequency domain frame signal; The front-stage generalized sidelobe canceling processor and the back-stage generalized sidelobe canceling processor are both configured to perform generalized sidelobe canceling processing on the array frequency domain frame signal output by the interleaving processor; the front-stage generalized sidelobe canceling processor and the back-stage generalized sidelobe canceling processor iteratively update equivalent weight vector information and transmit the equivalent weight vector information to each other; The signal fusion processor is configured to fuse beams output by all the back-stage generalized sidelobe canceling processors; The signal enhancement processor is configured to enhance the signal output by the signal fusion processor.
7. The apparatus of claim 6, wherein, The front-stage generalized sidelobe canceling processor and the back-stage generalized sidelobe canceling processor both include a main path and an auxiliary path; the main path includes a fixed beamformer, and the auxiliary path includes a blocking matrix and an adaptive noise canceling device; the blocking matrix is composed of an interference filter.
Citation Information
Patent Citations
Signal enhancement method based on microphone array
CN109389991A
Speech enhancement method based on generalized sidelobe cancellation structure
CN113362846A