Robust speech enhancement method based on adaptive beam forming and sparse spectrum constraint
By constructing a robust speech enhancement method with a generalized sidelobe canceller structure and sparse spectrum constraints, the problem of poor speech enhancement effect under strong interference and reverberation conditions is solved, and speech quality improvement and interference suppression are achieved in complex acoustic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-03
AI Technical Summary
Under conditions of strong interference and reverberation, the speech enhancement effect of existing adaptive beamforming methods such as the LS-GSC algorithm is significantly reduced, making it difficult to effectively improve the quality, clarity, and intelligibility of the target speech.
A robust speech enhancement method based on adaptive beamforming and sparse spectrum constraints is adopted. By constructing a signal model with a generalized sidelobe canceller structure, and combining the least squares criterion and the sparse speech spectrum constraint, the Lagrange multiplier alternating direction method is used for iterative solution to optimize the filter weight vector of the beamformer and output the enhanced speech signal.
It significantly improves speech quality under strong interference and reverberation environments, enhances the clarity and intelligibility of speech signals, improves the ability to suppress interference and reverberation, and exhibits stronger robustness.
Smart Images

Figure CN121789702A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech enhancement technology, and particularly relates to a robust speech enhancement method based on adaptive beamforming and sparse spectrum constraints. Background Technology
[0002] In acoustic applications such as hands-free voice communication and human-computer voice interaction, the speech signals captured by microphones are often affected by various interferences and reverberation, severely degrading the quality, clarity, and intelligibility of the target speech. Enhancing the target speech signal under conditions of strong interference and reverberation remains a challenging problem. In recent years, researchers have developed numerous single-channel and multi-channel speech enhancement methods. Compared to single-channel techniques, multi-channel methods, especially microphone array-based beamforming methods, can fully utilize spatial information to achieve more effective reverberation and interference suppression.
[0003] Typical microphone array beamforming methods include fixed beamforming and adaptive beamforming. Adaptive beamforming uses iterative algorithms to dynamically adjust array weights to adapt to time-varying acoustic environments, making it well-suited for scenarios with complex interference and reverberation. Commonly used adaptive beamforming methods include Minimum Variance Distortionless Response (MVDR) and Linear Constrained Minimum Variance (LCMV). MVDR beamformers aim to maintain a distortion-free response to the signal in the target direction while minimizing the output variance; while LCMV beamformers, based on MVDR, significantly suppress interference signals and improve beamformer performance by introducing zero-response constraints in multiple interference directions. The Generalized Sidelobe Canceller (GSC) reformulates the LCMV beamformer from a constrained optimization problem into an unconstrained optimization problem, providing implementation advantages while maintaining comparable performance. By stacking the output signals of the GSC into a vector, a GSC algorithm based on the least squares criterion (LS-GSC) can be established. Although some improved GSC variants utilize prior information to enhance beamformer performance, their performance still degrades significantly under severe interference and reverberation conditions. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a robust speech enhancement method based on adaptive beamforming and sparse spectral constraints, thereby resolving the issues present in the prior art.
[0005] To achieve the above objectives, this invention provides a robust speech enhancement method based on adaptive beamforming and sparse spectral constraints, comprising: Multi-channel observation signals are received through a microphone array, and a signal model of a generalized sidelobe canceller structure is constructed based on the multi-channel observation signals. Based on the signal model of the generalized sidelobe canceller structure, and combined with the least squares criterion and speech spectrum sparsity constraint, a beamforming optimization model is constructed. The beamforming optimization model is iteratively solved using the alternating direction method of Lagrange multipliers to obtain the adaptive filter weight vector of the generalized sidelobe canceller structure. Based on the adaptive filter weight vector, the enhanced speech signal is calculated and output.
[0006] Optionally, the process of constructing the signal model of the generalized sidelobe canceller structure includes: The filter weight vector of the beamformer is modeled as a linear combination of the fixed beamformer, the blocking matrix and the adaptive filter weight vector, and the filter structure of the generalized sidelobe canceller is determined. The multi-channel observation signal is filtered based on the blocking matrix to generate an input signal matrix; The multi-channel observation signal is filtered using the fixed beamformer to generate a reference signal. The difference between the reference signal and the product of the input signal matrix and the adaptive filter weight vector is used as the output signal of the beamformer.
[0007] Optionally, the process of constructing a beamforming optimization model includes: Based on the signal model of the generalized sidelobe canceller structure, the least squares error term is established by applying the least squares criterion; based on the short-time Fourier transform coefficients of the beamformer output signal, the sparse regularization term is established by applying the speech spectrum sparsity constraint; the least squares error term and the speech spectrum sparse regularization term are combined to establish the beamforming optimization model.
[0008] Optionally, the process of iteratively solving the beamforming optimization model using the Lagrange multiplier alternating direction method includes: By introducing auxiliary variables, the beamforming optimization model is transformed into a constrained optimization problem; Construct the augmented Lagrangian function corresponding to the constrained optimization problem; The solution is obtained iteratively by alternately minimizing the augmented Lagrange function and updating the variables until convergence is achieved.
[0009] Optionally, the process of alternately minimizing the augmented Lagrange function and updating the variables includes: The adaptive filter weight vector is updated by fixing the beamformer output signal and the Lagrange multiplier variables; the beamformer output signal is updated by fixing the adaptive filter weight vector, the Lagrange multiplier variables, and the auxiliary variables; the auxiliary variables are updated by fixing the Lagrange multiplier variables and the beamformer output signal; the Lagrange multiplier variables are updated based on the updated beamformer output signal and the adaptive filter weight vector; and the Lagrange multiplier variables are updated based on the updated beamformer output signal and the auxiliary variables.
[0010] Optionally, the iterative solution process further includes an initialization step: initializing the adaptive filter weight vector, the auxiliary variable, and the Lagrange multiplier variable as vectors, where the first element is 1 and the remaining elements are 0; initializing the beamformer output signal as a reference signal; and setting sparse regularization parameters and penalty parameters.
[0011] Optionally, the process of updating the adaptive filter weight vector includes: constructing a least squares problem based on the input signal matrix, the reference signal, the beamformer output signal, and the Lagrange multiplier variables, and obtaining the updated adaptive filter weight vector by solving the solution of the least squares problem.
[0012] Optionally, the process of updating the beamformer output signal includes: Based on the input signal matrix, the reference signal, the updated adaptive filter weight vector, the auxiliary variables, and the Lagrange multiplier variables, the updated beamformer output signal is obtained by minimizing the mean square error term.
[0013] Optionally, the process of updating auxiliary variables includes: The sum of the short-time Fourier transform coefficients of the beamformer output signal and the Lagrange multiplier variables updated in the previous iteration is calculated, and a soft thresholding operation is applied to each element of the sum.
[0014] Optionally, the process of updating the Lagrange multiplier variables includes: The difference between the updated auxiliary variable and the short-time Fourier transform coefficient of the beamformer output signal is calculated and added to the Lagrange multiplier variable of the previous round to obtain the updated Lagrange multiplier variable value.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: This invention proposes a robust beamforming method to improve speech quality in environments with strong interference and reverberation. Based on the traditional LS-GSC algorithm, this invention defines a speech spectrum sparsity penalty term according to the output signal of the GSC beamformer, establishing a novel beamforming optimization model. This sparse spectrum constraint enhances the robustness of the beamformer to strong interference and reverberation. The proposed beamforming optimization model is solved using the Alternating Direction Method of Lagrange Multipliers (ADMM) to obtain a beamforming algorithm with sparse spectrum constraints. Simulation experiments in a real acoustic environment further verify the effectiveness of this invention in suppressing strong reverberation and interference. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where speech signals act as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 3 For the reverberation time of this embodiment of the invention Furthermore, in an acoustic environment where fan noise serves as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 4 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where babble noise acts as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 5 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where speech signals act as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 6 For the reverberation time of this embodiment of the invention Furthermore, in an acoustic environment where fan noise serves as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 7 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where babble noise acts as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 8 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where speech signals act as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 9 For the reverberation time of this embodiment of the invention Furthermore, in an acoustic environment where fan noise serves as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 10 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where babble noise acts as interference, the LS-GSC algorithm and the proposed... - Comparison of the performance of the ADMM-GSC algorithm as a function of iSIR; Figure 11 For the reverberation time of this embodiment of the invention Furthermore, in acoustic environments where speech signals act as interference, the LS-GSC algorithm and the proposed... - Comparison of oSIR and iSIR for the ADMM-GSC algorithm. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] Example 1 like Figure 1 As shown, this embodiment provides a robust speech enhancement method based on adaptive beamforming and sparse spectral constraints, including: Multi-channel observation signals are received through a microphone array, and a signal model of a generalized sidelobe canceller structure is constructed based on the multi-channel observation signals. Based on the signal model of the generalized sidelobe canceller structure, and combined with the least squares criterion and speech spectrum sparsity constraint, a beamforming optimization model is constructed. The beamforming optimization model is iteratively solved using the alternating direction method of Lagrange multipliers to obtain the adaptive filter weight vector of the generalized sidelobe canceller structure. Based on the adaptive filter weight vector, the enhanced speech signal is calculated and output.
[0020] (I) Signal Model and Related Work: (1) Signal model: Considering a reverberant environment, this embodiment uses A linear array of equidistant omnidirectional microphones is used to pick up sound signals, with the distance between two adjacent microphones being [missing information]. Assuming there exists One sound source, front One is the expected source, the rest are... One is the source of interference. Without loss of generality, this embodiment uses the first microphone as the reference point, then in Time of the first The signal received by the microphone alone can be represented as: (1) In the formula, This represents the convolution operation. It is the first One source signal, From the first From the source to the first The impulse response of the microphone's acoustic channel only. It is the first Additive noise from only the microphone. (2) yes The vector form with length is , (3) yes The vector form with length is , This indicates the transpose operation.
[0021] If The recent If the samples are stacked into a vector, then (1) can be re-expressed as the following vector / matrix form: (4) In the formula, It is a length of microphone signal vector, It is a length of The source signal vector, , It is a length of The noise signal vector. It is a size The Sylvester matrix is specifically represented as follows: (5) The process of constructing the signal model of the generalized sidelobe canceller structure includes: The filter weight vector of the beamformer is modeled as a linear combination of the fixed beamformer, the blocking matrix and the adaptive filter weight vector, and the filter structure of the generalized sidelobe canceller is determined. The multi-channel observation signal is filtered based on the blocking matrix to generate an input signal matrix; The multi-channel observation signal is filtered using the fixed beamformer to generate a reference signal. The difference between the reference signal and the product of the input signal matrix and the adaptive filter weight vector is used as the output signal of the beamformer.
[0022] Beamforming aims to design a spatial filter that allows the desired source signal to be obtained from the output of the beamformer. To achieve this, the microphone signal... After filtering, the following output signal can be obtained: (6) In the formula, It is a length The filter weight vector, It is a length of The weight vector of the cascaded filter. It is a length of The noise signal vector, It is a size of The multi-channel convolution matrix is specifically represented as follows: (7) It is a by Composed of multi-channel convolutional matrices A 3D matrix is specifically represented as follows: (8) It is a length of The source signal vector is specifically represented as follows: (9) On the other hand, (6) can also be expressed as: (10) In the formula, It is a length of The observed signal vector.
[0023] (2) Related work: The LCMV algorithm is a commonly used speech enhancement method designed to extract the desired speech signal from array observation signals. Under the condition of satisfying several linear constraints, this algorithm aims to enhance the signal in the desired direction while suppressing interference and noise from other directions by minimizing the power of the microphone array beamformer output signal. The optimization model of the LCMV method is as follows: (11) In the formula This represents taking the mathematical expectation. It is a length of The gain vector is defined as: (12) In the formula, It is a length of The zero vector, It is by A length of The vector consisting of unit vectors has a length of ,in It is a length of unit vector, subscript This represents the position of 1 in a unit vector, and... The index of the largest diagonal element corresponds to this.
[0024] A common approach to implementing the LCMV algorithm is to reformulate the constrained optimization problem as an unconstrained optimization problem using a Gaussian Filter (GSC) structure. Within this framework, the beamforming filter vector... The signal is decomposed into two sub-filters in mutually orthogonal subspaces. The first sub-filter acts as a fixed beamformer, designed to enhance the speech signal in the target direction while attenuating interference and noise from other directions. The second sub-filter consists of a blocking matrix used to suppress the target speech signal, allowing only interference and noise components to pass through. The output is then adaptively filtered to estimate and eliminate residual interference and noise leaking through the fixed beamformer. The beamformer's output signal is obtained by subtracting the adaptive path's output signal from the fixed beamformer's output signal. Based on the GSC principle, this embodiment defines a size of... Blocking matrix Its column vectors are matrices The basis of the orthogonal complement space spanned by the column vectors of . According to the definition of the orthogonal complement space, we have Therefore, the following options are available: (13) In the formula, It is the size of unit array, It is the size of The zero matrix.
[0025] Based on the blocking matrix Define a size of matrix and a length of vector ,in, It is a length of column vectors, It is a length of And the weight vector is unaffected by constraints. According to the above definition, the filter weight vector... It can be represented as ,Right now: (14) In the formula, This represents a fixed beamformer that satisfies linear constraints. Based on the decomposition of (14) and the constraints in (11), we can obtain: (15) The solution can be obtained Therefore, the solution for a fixed beamformer can be obtained as follows: (16) Based on (10), dimensional output signal vector It can be represented as (17) In the formula, It's a microphone signal. Multi-channel matrix yes The microphone signal matrix. Substituting (14) into (17), we get: (18) In the formula, It is a length of The reference signal vector, It is the size of The input signal matrix, It is a length of The filter vector to be estimated. Filter The estimate can be obtained online using standard adaptive algorithms, such as Least Mean Square (LMS) or Recursive Least Squares (RLS) adaptive algorithms, or offline using statistical methods, such as Wiener filters or LS criteria. This embodiment aims to utilize the spectral sparsity of speech signals to construct a robust optimization model based on a GSC beamformer and solve it offline, thereby obtaining a robust speech enhancement method.
[0026] The process of constructing a beamforming optimization model includes: Based on the signal model of the generalized sidelobe canceller structure, the least squares error term is established by applying the least squares criterion; based on the short-time Fourier transform coefficients of the beamformer output signal, the sparse regularization term is established by applying the speech spectrum sparsity constraint; the least squares error term and the speech spectrum sparse regularization term are combined to establish the beamforming optimization model.
[0027] (II) Optimization Model and Solution: According to (18), the filter vector This can be achieved by minimizing the cost function. Direct solution, where express Norm. Its algorithm corresponds to the standard LS-GSC algorithm. This embodiment considers introducing spectral sparsity constraints on the beamformer output speech signal within the LS-GSC framework to enhance the robustness of the GSC beamformer to strong interference and reverberation.
[0028] (1) Optimization model: Extensive experiments have shown that clean speech signals exhibit strong sparsity in the Short Time Fourier Transform (STFT) domain. However, when speech is contaminated by reverberation, the sparsity of the microphone speech signal is significantly reduced. By forcibly constraining the sparsity of the beamformer output speech signal in the optimization model, enhanced speech can be obtained at the beamformer output even under conditions of strong interference and reverberation. Therefore, this embodiment establishes the following optimization model: (19) In the formula, express Norm, It is a sparsity regularization parameter used to strike a trade-off between minimizing the LS error and the sparsity of the STFT coefficients of the output signal. It is the STFT operator, which will 3D time-domain vector Transform into A time-frequency domain vector. Specifically, the STFT coefficients of the beamformer output signal are calculated as follows: (20) In the formula, It is a frequency band index. It's the number of bandwidths. It is a frame index. For frame number, It is the STFT analysis window. It's a frame shift. Define the STFT operator. , , It is composed of all STFT coefficients Composition. When using a compressed STFT analysis window, the inverse STFT operator satisfy For simplicity, the remaining parts of this embodiment will omit the time variable. .
[0029] (2) Solution of the optimization model: This embodiment uses the ADMM method to solve the optimization problem in (19).
[0030] The process of iteratively solving the beamforming optimization model using the alternating direction method with Lagrange multipliers includes: By introducing auxiliary variables, the beamforming optimization model is transformed into a constrained optimization problem; Construct the augmented Lagrangian function corresponding to the constrained optimization problem; The solution is obtained iteratively by alternately minimizing the augmented Lagrange function and updating the variables until convergence is achieved.
[0031] By introducing an auxiliary variable The optimization problem in (19) can be restated as: (twenty one) Therefore, the augmented Lagrangian function corresponding to (21) is as follows: (twenty two) In the formula, For length is Lagrange multiplier vectors, It is another length of Lagrange multiplier vectors, It is a penalty parameter.
[0032] The process of alternately minimizing the augmented Lagrange function and updating the variables includes: The adaptive filter weight vector is updated by fixing the beamformer output signal and the Lagrange multiplier variables; the beamformer output signal is updated by fixing the adaptive filter weight vector, the Lagrange multiplier variables, and the auxiliary variables; the auxiliary variables are updated by fixing the Lagrange multiplier variables and the beamformer output signal; the Lagrange multiplier variables are updated based on the updated beamformer output signal and the adaptive filter weight vector; and the Lagrange multiplier variables are updated based on the updated beamformer output signal and the auxiliary variables.
[0033] (22) respectively apply to the variables , and Alternate minimization, and consider augmented Lagrange multiplier vectors. and The following optimization subproblems and update equations can be obtained: (twenty three) (twenty four) (25) (26) (27) In the formula, superscript ( The number of iterations is represented by . Therefore, these subproblems are solved alternately until the termination criterion is met or the maximum number of iterations is reached. The solutions to the first three subproblems are derived below.
[0034] First, the update rule for the filter vector can be obtained by minimizing the objective function in (23), i.e. (28) The process of updating the adaptive filter weight vector includes: constructing a least squares problem based on the input signal matrix, the reference signal, the beamformer output signal, and the Lagrange multiplier variables; and obtaining the updated adaptive filter weight vector by solving the solution of the least squares problem.
[0035] Secondly, minimizing (24) yields the update rule for the beamformer output vector, i.e. (29) The process of updating the beamformer output signal includes: Based on the input signal matrix, the reference signal, the updated adaptive filter weight vector, the auxiliary variables, and the Lagrange multiplier variables, the updated beamformer output signal is obtained by minimizing the mean square error term.
[0036] Finally, the solution to (25) corresponds to the proximal mapping of the penalty function. For The norm indicates that the mapping has a closed-form solution. To simplify the notation, an intermediate variable is defined. . The near-end mapping of the norm is given by element-wise soft thresholding, i.e.: (30) In the formula, The vector represents the first One element, .
[0037] The process of updating auxiliary variables includes: Calculate the sum between the short-time Fourier transform coefficients of the beamformer output signal and the Lagrange multiplier variables updated in the previous iteration, and apply a soft thresholding operation to each element of the sum.
[0038] The difference between the updated auxiliary variable and the short-time Fourier transform coefficient of the beamformer output signal is calculated and added to the Lagrange multiplier variable of the previous round to obtain the updated Lagrange multiplier variable value.
[0039] The filter vector can be effectively solved by iteratively calculating (28), (29), (30) and (26) and (27). Then, based on (14) and (16), the global GSC beamforming filter vector can be obtained. Because the proposed method introduces Norm constraints are used, and the optimization problem is solved using the ADMM framework, hence the name. -ADMM-GSC algorithm. The iterative solution process also includes an initialization step: the adaptive filter weight vector, the auxiliary variable, and the Lagrange multiplier variable are initialized as a vector, where the first element of the vector is 1 and the remaining elements are 0. The beamformer output signal is initialized as a reference signal, and sparse regularization parameters and penalty parameters are set.
[0040] (III) Experimental verification; This embodiment compares the LS-GSC algorithm and the proposed algorithm through experimental research. The performance of the ADMM-GSC algorithm. For the algorithm proposed in this invention, the parameter vector... , and All initialized to ,vector Initialize as a vector The value of the filter length Impulse response length Frame length Frame count .
[0041] Microphone signals from experiments were simulated using a multi-channel acoustic impulse response database, which consists of real room acoustic impulse responses measured by the Speech and Acoustics Laboratory at Bar-Ilan University. It includes three reverberation environments: mild, moderate, and severe, with corresponding reverberation times of [missing information]. , and The experiments in this embodiment verify the performance of the proposed algorithm under these three reverberation conditions. The laboratory size is... The acoustic impulse response sampling rate was 16 kHz. A uniform linear array of six omnidirectional microphones was used to pick up the acoustic signal, with an adjacent microphone spacing of 0.08 m. In the experiment, a target source was considered ( ) and an interference source ( The two sound sources are located at 45° and 15° azimuths, respectively, with a distance of 1m from the center of the microphone array. The microphone signal is obtained by convolving the sound source signal with its corresponding room acoustic impulse response. The target speech signal is taken from the TIMIT database. This embodiment's experiments verified three types of interference: the first is a speech signal taken from the Librispeech database, the second is a recorded fan noise, and the third is a recorded bubble noise. The duration of each sound source signal is 8s. To focus on verifying the effectiveness of the proposed algorithm, additive noise was not included in the experiments.
[0042] To evaluate the speech enhancement performance of the proposed algorithm, two widely used metrics were employed: Perceptual Speech Quality Assessment (PESQ) and Cepstral Distance (CD). By definition, a higher PESQ score indicates better speech quality, while a lower CD value indicates less speech distortion.
[0043] To quantitatively evaluate the interference suppression performance of the algorithm, the output signal-to-interference ratio (oSIR) is used as the performance metric. The input signal-to-interference ratio (iSIR) is defined as follows: (31) In the formula, It is the first A desired signal source, It is the first One source of interference signal. oSIR is defined as: (32) In the formula, From the first The desired signal source to the first The impulse response of the microphone's acoustic channel only. From the first The interference signal source to the first This refers to the impulse response of the microphone's acoustic channel. If the output signal-to-interference ratio (oSIR) is greater than the input signal-to-interference ratio (iSIR), it means that the algorithm's interference suppression performance is improved. The larger the difference between the two, the better the interference suppression effect.
[0044] Figure 2 Demonstrated in reverberation time Furthermore, in acoustic environments where speech serves as interference, the LS-GSC algorithm and the proposed... The PESQ and CD values of the ADMM-GSC algorithm vary with iSIR. Figure 2 As can be seen in (a), the PESQ values of both algorithms increase with increasing iSIR. However, the proposed... The ADMM-GSC algorithm consistently outperforms the LS-GSC algorithm, indicating that the former has stronger robustness against reverberation and interference. Specifically, the proposed algorithm demonstrates stronger interference suppression capability at low iSIR and stronger dereverberation performance at high iSIR. Figure 2 As can be seen from (b) above, the CD value of the proposed algorithm is consistently lower than that of the traditional LS-GSC algorithm, especially under high iSIR conditions. This indicates that, compared to the traditional LS-GSC algorithm, the proposed algorithm... The ADMM-GSC algorithm outputs speech with lower distortion because the latter's optimization model introduces the sparsity of the speech amplitude spectrum. Figure 3 and Figure 4 Demonstrated in reverberation time In an acoustic environment where fan noise and bubble noise are considered as interferences, the LS-GSC algorithm and the proposed... The PESQ and CD values of the ADMM-GSC algorithm vary with iSIR. Figure 3 and Figure 4 All of these reflect the points raised. The ADMM-GSC algorithm significantly outperforms the traditional LS-GSC algorithm in two performance metrics, indicating that the proposed algorithm has a stronger speech enhancement effect in low reverberation environments without speech interference.
[0045] Figure 5 , Figure 6 and Figure 7 Demonstrated in reverberation time Furthermore, in acoustic environments where speech, fan noise, and babble noise are respectively considered as interferences, the LS-GSC algorithm and the proposed... The PESQ and CD values of the ADMM-GSC algorithm vary with iSIR. Figure 5 , Figure 6 and Figure 7 It can be seen that when the reverberation time At that time, what was mentioned The ADMM-GSC algorithm outperforms the traditional LS-GSC algorithm in both PESQ and CD metrics in environments with and without speech interference, indicating that the proposed algorithm also has a stronger speech enhancement effect in moderate reverberation environments. Figure 8 , Figure 9 and Figure 10 Demonstrated in reverberation time Furthermore, in acoustic environments where speech, fan noise, and babble noise are respectively considered as interferences, the LS-GSC algorithm and the proposed... The PESQ and CD values of the ADMM-GSC algorithm vary with iSIR. Figure 8 , Figure 9 and Figure 10 It can be seen that when the reverberation time In practice, the proposed algorithm achieves better interference suppression and speech dereverberation capabilities than the LS-GSC algorithm. These experimental results validate the effectiveness of the proposed optimization model and beamforming algorithm under reverberation and interference environments of varying intensities.
[0046] Figure 11 Demonstrated in reverberation time Furthermore, in acoustic environments where speech serves as interference, the LS-GSC algorithm and the proposed... The relationship between oSIR and iSIR in the ADMM-GSC algorithm. It can be seen that the oSIR of both algorithms increases with increasing iSIR, but the proposed... The oSIR value of the ADMM-GSC algorithm is consistently significantly higher than that of the LS-GSC algorithm, meaning that regardless of the strength of the interference, the proposed... The ADMM-GSC algorithm outperforms the LS-GSC algorithm in suppressing interference and reverberation, further validating the robustness of the proposed algorithm in speech enhancement.
[0047] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A robust speech enhancement method based on adaptive beamforming and sparse spectral constraints, characterized in that, Includes the following steps: Multi-channel observation signals are received through a microphone array, and a signal model of a generalized sidelobe canceller structure is constructed based on the multi-channel observation signals. Based on the signal model of the generalized sidelobe canceller structure, and combined with the least squares criterion and speech spectrum sparsity constraint, a beamforming optimization model is constructed. The beamforming optimization model is iteratively solved using the alternating direction method of Lagrange multipliers to obtain the adaptive filter weight vector of the generalized sidelobe canceller structure. Based on the adaptive filter weight vector, the enhanced speech signal is calculated and output.
2. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 1, characterized in that, The process of constructing the signal model of the generalized sidelobe canceller structure includes: The filter weight vector of the beamformer is modeled as a linear combination of the fixed beamformer, the blocking matrix and the adaptive filter weight vector, and the filter structure of the generalized sidelobe canceller is determined. The multi-channel observation signal is filtered based on the blocking matrix to generate an input signal matrix; The multi-channel observation signal is filtered using the fixed beamformer to generate a reference signal. The difference between the reference signal and the product of the input signal matrix and the adaptive filter weight vector is used as the output signal of the beamformer.
3. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 2, characterized in that, The process of constructing a beamforming optimization model includes: Based on the signal model of the generalized sidelobe canceller structure, the least squares error term is established by applying the least squares criterion; based on the short-time Fourier transform coefficients of the beamformer output signal, the sparse regularization term is established by applying the speech spectrum sparsity constraint; the least squares error term and the speech spectrum sparse regularization term are combined to establish the beamforming optimization model.
4. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 3, characterized in that, The process of iteratively solving the beamforming optimization model using the alternating direction method with Lagrange multipliers includes: By introducing auxiliary variables, the beamforming optimization model is transformed into a constrained optimization problem; Construct the augmented Lagrangian function corresponding to the constrained optimization problem; The solution is obtained iteratively by alternately minimizing the augmented Lagrange function and updating the variables until convergence is achieved.
5. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 4, characterized in that, The process of alternately minimizing the augmented Lagrange function and updating the variables includes: The adaptive filter weight vector is updated by fixing the beamformer output signal and the Lagrange multiplier variables; the beamformer output signal is updated by fixing the adaptive filter weight vector, the Lagrange multiplier variables, and the auxiliary variables; the auxiliary variables are updated by fixing the Lagrange multiplier variables and the beamformer output signal; the Lagrange multiplier variables are updated based on the updated beamformer output signal and the adaptive filter weight vector; and the Lagrange multiplier variables are updated based on the updated beamformer output signal and the auxiliary variables.
6. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 5, characterized in that, The iterative solution process also includes an initialization step: initializing the adaptive filter weight vector, the auxiliary variable, and the Lagrange multiplier variable as vectors, where the first element is 1 and the remaining elements are 0; initializing the beamformer output signal as a reference signal, and setting sparse regularization parameters and penalty parameters.
7. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 6, characterized in that, The process of updating the adaptive filter weight vector includes: constructing a least squares problem based on the input signal matrix, the reference signal, the beamformer output signal, and the Lagrange multiplier variables; and obtaining the updated adaptive filter weight vector by solving the solution of the least squares problem.
8. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 6, characterized in that, The process of updating the beamformer output signal includes: Based on the input signal matrix, the reference signal, the updated adaptive filter weight vector, the auxiliary variables, and the Lagrange multiplier variables, the updated beamformer output signal is obtained by minimizing the mean square error term.
9. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 6, characterized in that, The process of updating auxiliary variables includes: The sum of the short-time Fourier transform coefficients of the beamformer output signal and the Lagrange multiplier variables updated in the previous iteration is calculated, and a soft thresholding operation is applied to each element of the sum.
10. The robust speech enhancement method based on adaptive beamforming and sparse spectral constraints according to claim 6, characterized in that, The process of updating the Lagrange multiplier variables includes: The difference between the updated auxiliary variable and the short-time Fourier transform coefficient of the beamformer output signal is calculated and added to the Lagrange multiplier variable of the previous round to obtain the updated Lagrange multiplier variable value.