A Low-Latency Speech Enhancement Method Based on Microphone Array

By using OBF filters and artificial neural networks to construct orthogonal basis function models in microphone arrays, the problems of large residual noise and time delay in existing speech enhancement methods are solved, achieving better speech enhancement performance and reduced network latency with short filter lengths.

CN119229885BActive Publication Date: 2025-11-14NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371287.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-11-14
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing deep learning-based speech enhancement methods suffer from high residual noise and long system delay in frequency domain beamforming, while time domain filtering and summing networks increase network delay when the filter length is long, making it difficult to achieve good speech enhancement performance with short filter lengths.

Method used

A low-latency speech enhancement method based on microphone arrays is adopted. By using OBF filters in each channel of the microphone array and combining them with artificial neural networks, an orthogonal basis function model and an adaptive beamformer are constructed to optimize pole parameters, reduce network latency, and improve speech enhancement performance.

Benefits of technology

It achieves better speech enhancement with short filter length, reduces network latency, and is applicable to various noise environments. It is versatile and flexible, and adapts to adaptive beamforming under different conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229885B_ABST
    Figure CN119229885B_ABST
Patent Text Reader

Abstract

This invention discloses a low-latency speech enhancement method based on a microphone array. The method includes: setting a set of initial pole parameters; optimizing the initial pole parameters using an artificial neural network to obtain real poles; constructing orthogonal basis function models for each channel of the microphone array using the real poles and calculating the responses of filters of each order; performing frame segmentation and temporal feature extraction on the received signal from the microphone array, and estimating the adaptive beamformer weights constructed from the orthogonal basis function models using an improved temporal network; calculating the system response of each channel of the beamforming network based on the filter responses and beamformer weights to obtain the enhanced complete speech signal. This invention, by using an orthogonal basis structure beamforming network, can flexibly adjust the poles, increase the network's degrees of freedom, shorten the filter length, and reduce network latency; achieving better speech enhancement results with shorter filter lengths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to speech processing technology, and in particular to a low-latency speech enhancement method based on a microphone array. Background Technology

[0002] Speech enhancement is a technique that eliminates background noise or directional interference during speech propagation to improve speech quality. Multichannel beamforming, in particular, can effectively reduce speech distortion by fully utilizing temporal and spatial information. Deep learning (DL)-based speech enhancement methods leverage the complex structure and strong expressive power of deep neural networks (DNNs) to learn the relationship between the microphone-received signal and the desired signal and their characteristics. This approach is applicable under various background noise conditions and is a highly efficient technique.

[0003] Existing frequency domain beamforming methods based on deep learning mainly estimate the covariance matrix of the signal and then use it in minimum variance distortionless response (MVDR) beamforming. These methods can improve speech quality to some extent, but they have large residual noise. In addition, in order to ensure the frequency domain resolution, the frame length is relatively long, and the contextual information of the signal needs to be extracted to enrich the features during the training process, so the system latency is also relatively large.

[0004] Existing time-domain filter-and-sum network (FaSNet) adaptively estimates beamformer weights in the time domain, which can process short-frame-length speech signals, ensuring real-time performance and better noise reduction. However, FaSNet's latency is related to the filter length, requiring a longer filter length to achieve good filtering results, thus increasing network latency. Summary of the Invention

[0005] Purpose of the invention: To address the above problems, the purpose of this invention is to provide a low-latency speech enhancement method based on a microphone array. By using OBF filters in each channel of the microphone array to increase the degree of freedom, good speech enhancement performance is achieved with a short filter length. The entire beamforming framework is implemented using an artificial neural network, while reducing network latency.

[0006] Technical solution: The present invention provides a low-latency speech enhancement method based on a microphone array, comprising the following steps:

[0007] Step 1: Set a set of initial pole parameters, optimize the initial pole parameters using an artificial neural network to obtain real poles, and use these real poles as parameters for the multi-pole orthogonal basis function model of each channel of the microphone array;

[0008] Step 2: Construct orthogonal basis function models for each channel of the microphone array using real poles, and calculate the responses of each order of filters;

[0009] Step 3: Frame the received signal from the microphone array and extract temporal features, and use an improved temporal network to estimate the weights of the adaptive beamformer constructed from the orthogonal basis function model.

[0010] Step 4: Calculate the system response of each channel of the beamforming network based on the filter response and beamformer weights, and filter and sum the received signal of the microphone array to obtain the enhanced frame-level speech signal. Perform an overlap and summation operation on the enhanced frame-level speech signal to obtain the enhanced complete speech signal.

[0011] Further, step 1 includes:

[0012] For a microphone array with m channels, m = 1, 2, ..., M, an unbiased one-dimensional convolutional neural network with 1 input channel and L output channels is used to train the initial pole parameters. M represents the number of microphones, L represents the filter length, and the kernel size and stride of the one-dimensional convolutional neural network are both set to 1. Data with a constant value of 1 is used as input, and the output is the learnable weights of this one-dimensional convolutional neural network, expressed as:

[0013]

[0014] In the formula, Conv1d represents a one-dimensional convolutional neural network; l represents the filter order number;

[0015] The extreme points obtained in each training session Perform boundary processing, and then process the pole a. ml As poles of the orthogonal basis function model; where the boundary treatment is expressed as:

[0016]

[0017] In the formula, HardTanh(·) is the activation function, expressed as:

[0018]

[0019] In the formula, a max and a min These represent the maximum and minimum values ​​of the linear range of the HardTanh(·) function, respectively.

[0020] Further, step 2 includes:

[0021] The state-space equation of the orthogonal basis function model of the m-th channel of the microphone array is:

[0022] x m (t)=A m x m (t-1)+b m u m (t)

[0023]

[0024] In the formula, Let be the state vector at the t-th discrete time point. The state matrix, For the input-state vector, u m (t) represents the signal received by the microphone, w m This refers to the state-output vector, i.e., the weights of the orthogonal basis function model. For the output of the orthogonal basis function model, [·] T Represents the transpose of a matrix or vector;

[0025] Calculate the state matrix A based on the state-space equation of the orthogonal basis function model. m The (j,l) elements and the input-state vector b m The l-th element is:

[0026]

[0027] in, This is the orthogonal state-space realization matrix of the l-th first-order all-pass filter that forms the orthogonal basis function model of the m-th channel, and the elements in the matrix are related to the poles a. ml The relationship is: d ml =-a ml ;

[0028] Set the initial state vector x m If (0) = 0, then the state-space equation of the orthogonal basis function model is:

[0029]

[0030] In the formula, i represents the discrete-time index. Represents the state matrix A m The i-1th power;

[0031] Therefore, the length is L. h The response of the L-order filter is:

[0032]

[0033] in, Let L represent the response vector of the L-th filter at time point i, where i = 1, 2, ..., L.h L h ≥L.

[0034] Furthermore, step 3 includes:

[0035] If the microphone received signal is divided into frames with a frame length of W and a frame shift of J = W / 2, then the k-th frame signal of the m-th channel of the microphone array is represented as:

[0036] u m (k)=[u m ((k-1)J+1),…,u m ((k-1)J+W)] T

[0037] In the case of non-causal feature extraction and filtering, concatenating the current frame signal and its context information yields the following concatenated signal:

[0038] v m (k)=[u m ((k-1)J-C+1),…,u m ((k-1)J+W+C)] T

[0039] In the formula, C represents the length of the preceding or following information;

[0040] Signal u is received via reference microphone ref (k) and v m (k) Extract the feature vector ξ of each channel of the microphone array m (k), and the eigenvector ξ m (k) is input into the improved temporal network SeqNet(·) to obtain the weights of the adaptive wave velocity former.

[0041] Furthermore, the improved temporal network SeqNet(·) begins with a linear bottleneck layer with B output channels. The weights of the frame-level adaptive beamformers corresponding to the state-space equations are then estimated as follows:

[0042] w m (k)=OutputLayer(SeqNet(ξ m (k)))

[0043] In the formula, w m (k)=[w m,1 (k),…,w m,L (k)] T This is the composition vector of the adaptive wave velocity generator weights;

[0044] The output layer (OutputLayer) of the improved temporal network SeqNet(·) consists of a one-dimensional convolutional neural network and an activation function, with the following specific settings:

[0045] OutputLayer(p)=PReLU(n(Wp+q))

[0046] In the formula, This is the output of a time-series network. and η represents the weights and biases of the one-dimensional convolutional neural network, η is the scale factor that controls the output of the one-dimensional convolutional neural network, and PReLU(·) is the activation function of the parameter-corrected linear unit.

[0047] Furthermore, in step 4, the system response of each channel of the beamforming network is as follows:

[0048]

[0049] When L h When the value is greater than L, pad with L before convolutional concatenation. h -L zeros are added to ensure that the length of the convolutional signal is equal to the expected signal length. The concatenated information after zero padding is then:

[0050]

[0051] A one-dimensional convolutional neural network is used to implement beamforming network filtering of the received signals of each channel, and the output signals of all channels are summed to obtain the enhanced k-th frame speech signal:

[0052]

[0053] In the formula, * represents the convolution operation of a one-dimensional convolutional neural network, and the relationship between the filter length and the context signal length is L = 2C + 1.

[0054] Beneficial effects: Compared with the prior art, the significant advantages of this invention are:

[0055] 1. In this invention, the FasNet with all poles set to zero is extended into an orthogonal basis structure beamforming network, which allows for flexible adjustment of poles and increases the network's degree of freedom.

[0056] 2. By training a one-dimensional convolutional neural network, the problem of finding multiple poles in the non-convex OBF model is solved. Unified orthogonal basis extension results are given for speech under different conditions. The pole solution and beamformer weight estimation are decoupled, making it more flexible and suitable for adaptive beamforming in actual situations.

[0057] 3. Compared with frequency domain methods, the present invention has better noise reduction effect; compared with time domain filtering and summation methods, the present invention can achieve better speech enhancement performance when the filter length is short.

[0058] 4. The speech enhancement method proposed in this invention is universal and is not affected by the type of temporal network module. It has a good improvement effect under different types of networks. Attached Figure Description

[0059] Figure 1 This is a flowchart of a low-latency speech enhancement method based on a microphone array;

[0060] Figure 2 This is a flowchart of a low-latency speech enhancement method based on a microphone array;

[0061] Figure 3 This is a schematic diagram showing the distribution of the indoor microphone array and sound source locations. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.

[0063] The low-latency speech enhancement method based on a microphone array described in this embodiment includes at least the following steps 1 to 4, as shown in the flowchart below. Figure 1 and Figure 2 As shown, this is a microphone array voice enhancement applicable to any array configuration, designed to suppress background noise in the microphone received signal using beamforming technology, without considering room reverberation removal.

[0064] In this embodiment, a microphone array consisting of M array elements is used to receive the voice signal. Therefore, the m-th microphone receives the noisy signal as follows:

[0065] u m (t)=y m (t)+n m (t), m=1,2,…,M

[0066] In the formula, The noise-free signal received by the microphone, i.e., the desired signal sought in this invention, g m (t) represents the room impulse response from the unknown location sound source s(t) to the m-th microphone. For convolution operators, n m (t) represents the additive noise at the m-th microphone. Here, it is assumed that the noise and the source signal are uncorrelated and that each additive noise is independent of the others.

[0067] In this example, a low-latency speech enhancement method based on a microphone array includes the following steps:

[0068] Step 1: Set a set of initial pole parameters, optimize the initial pole parameters using an artificial neural network to obtain real poles, and use these real poles as parameters for the multi-pole orthogonal basis function model of each channel of the microphone array;

[0069] Step 2: Construct orthogonal basis function models for each channel of the microphone array using real poles, and calculate the responses of each order of filters;

[0070] Step 3: Frame the received signal from the microphone array and extract temporal features, and use an improved temporal network to estimate the weights of the adaptive beamformer constructed from the orthogonal basis function model.

[0071] Step 4: Calculate the system response of each channel of the beamforming network based on the filter response and beamformer weights, and filter and sum the received signals from the microphone array to obtain the enhanced frame-level speech signal. Perform an overlap-addition operation on the enhanced frame-level speech signal to obtain the enhanced complete speech signal.

[0072] Further, step 1 includes:

[0073] A microphone array consisting of M elements is used to receive signals. Different orthonormal basis function (OBF) models (i.e., OBF filters) are applied to different microphone channels. For each microphone array channel m, m = 1, 2, ..., M, an unbiased one-dimensional convolutional neural network with 1 input channel and L output channels is used to train the initial pole parameters. M represents the number of microphones, L represents the filter length, and the kernel size and stride of the one-dimensional convolutional neural network are set to 1. Data with a constant value of 1 is used as input, and the output is the learnable weights of this one-dimensional convolutional neural network. These weights are used as the pole parameters of the orthogonal basis function model. The training process is represented as follows:

[0074]

[0075] In the formula, Convld represents a one-dimensional convolutional neural network; l represents the filter order number;

[0076] The initial pole parameters obtained from each training session Boundary treatment is performed to ensure system stability, and the poles a after boundary treatment are... ml As poles of the orthogonal basis function model; where the boundary treatment is expressed as:

[0077]

[0078] In the formula, HardTanh(·) is the activation function, expressed as:

[0079]

[0080] In the formula, a max and a min These represent the maximum and minimum values ​​of the linear range of the HardTanh(·) function, respectively. In this example, the linear range can be set to [0, 0.9999], i.e., a min Corresponding to 0, a max The corresponding value is 0.9999.

[0081] Further, step 2 includes:

[0082] The state-space equation of the orthogonal basis function model of the m-th channel of the microphone array is expressed as:

[0083] x m (t)=A m x m (t-1)+b m u m (t)

[0084]

[0085] In the formula, Let be the state vector at t discrete time points. The state matrix, For the input-state vector, u m (t) represents the signal received by the microphone, w m This refers to the state-output vector, i.e., the weights of the orthogonal basis function model. For the output of the orthogonal basis function model, [·] T Represents the transpose of a matrix or vector;

[0086] Based on the state-space equation of the orthogonal basis function model and the real pole a ml Calculate state matrix A m The (j,l) elements and the input-state vector b m The l-th element is represented as:

[0087]

[0088] in, This is the orthogonal state-space realization matrix of the l-th first-order all-pass filter that forms the orthogonal basis function model of the m-th channel, and the elements in the matrix are related to the poles a. ml The relationship is: d ml =-a ml ;

[0089] Set the initial state vector x mIf (0) = 0, then the state-space equation of the orthogonal basis function model is expressed as:

[0090]

[0091] In the formula, i represents the discrete-time index. Represents the state matrix A m The i-1th power;

[0092] Through state matrix A m and input-state vector b m The calculated length is L. h The response of the L-order filter is:

[0093] f m =[f mLh ,…,f mi ,…,f m2 ,f m1 ]

[0094] in, Let L represent the response vector of the L-th filter at time point i, where i = 1, 2, ..., L. h L h ≥L.

[0095] Furthermore, step 3 includes:

[0096] If the microphone received signal is divided into frames with a frame length of W and a frame shift of J = W / 2, then the k-th frame signal of the m-th channel of the microphone array is represented as:

[0097] u m (k)=[u m ((k-1)J+1),…,u m ((k-1)J+W)] T

[0098] In the case of non-causal feature extraction and filtering, concatenating the current frame signal and its context information yields the following concatenated signal:

[0099] v m (k)=[u m ((k-1)J-C+1),…,u m ((k-1)J+W+C)] T

[0100] In the formula, C represents the length of the preceding or following information;

[0101] Signal u is received via reference microphone ref (k) and v m (k) Extract the feature vector ξ of each channel of the microphone array m(k), and the eigenvector ξ m (k) is input into the improved temporal network SeqNet(·) to obtain the weights of the adaptive wave velocity former.

[0102] In one example, the microphone received signal is framed. For instance, if the frame length W = 80 (i.e., 10ms per frame) and the frame shift H = W / 2, the k-th frame signal of the m-th channel can be obtained. If microphone 1 is selected as the reference channel, under non-causal feature extraction and filtering conditions, the current frame signal of each channel and its context information are spliced ​​together to obtain the spliced ​​information, and the feature vector ξ of each channel input to the temporal network is calculated. m (k).

[0103] A temporal network, SeqNet(·), is constructed by stacking Temporal Convolutional Network (TCN) or Dual-Path Recurrent Neural Network (DPRNN) modules. This temporal network is then improved by stacking a one-dimensional convolutional neural network and an activation function layer after the output layer of SeqNet(·). Specifically, if the TCN network is directly configured for real-time online data processing, the network latency using the TCN module is W+C. If the DPRNN network is modified for real-time online data processing, with intra-block RNNs using bidirectional long short-term memory units and inter-block RNNs using unidirectional long short-term memory units, and the normalization layer using causal normalization, the network latency using the DPRNN module is (K+1)W / 2+C, where K=4 is the block length of the temporal sequence processed by the DPRNN.

[0104] The improved temporal network SeqNet(·) begins with a linear bottleneck layer of B output channels. Therefore, the weights of the frame-level adaptive beamformers corresponding to the state-space equations are estimated as follows:

[0105] w m (k)=OutputLayer(SeqNet(ξ m (k)))

[0106] In the formula, w m (k)=[w m,1 (k),…,w m,L (k)] T This is the composition vector of the adaptive wave velocity generator weights;

[0107] The output layer (OutputLayer) of the improved temporal network SeqNet(·) consists of a one-dimensional convolutional neural network and an activation function, with the following specific settings:

[0108] OutputLayer(p)=PReLU(n(Wp+q))

[0109] In the formula, This is the output of a time-series network. and η represents the weights and biases of the one-dimensional convolutional neural network, η is the scale factor that controls the output of the one-dimensional convolutional neural network, and PReLU(·) is the activation function of the parameter-corrected linear unit.

[0110] Furthermore, by calculating the responses of each order filter and the adaptive beamformer weight vector, the system response in each channel of the beamforming network is obtained as follows:

[0111]

[0112] When L h When the value is greater than L, pad with L before convolutional concatenation. h -L zeros are added to ensure that the length of the convolutional signal is equal to the expected signal length. The concatenated information after zero padding is then:

[0113]

[0114] A one-dimensional convolutional neural network is used to implement beamforming network filtering of the received signals of each channel, and the output signals of all channels are summed to obtain the enhanced k-th frame speech signal:

[0115]

[0116] In the formula, * represents the convolution operation of a one-dimensional convolutional neural network, and the relationship between the filter length and the context signal length is L = 2C + 1.

[0117] In this example, the scale-invariant source-to-noise ratio (SI-SNR) is used as the loss function. An Adam optimizer with an initial learning rate of 0.001 is used to train the beamforming network proposed in this invention. The improvement in perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI), as well as SI-SNR and signal-to-distortion ratio (SDR) are used as performance evaluation metrics.

[0118] The invention will be explained in detail below with simulation examples.

[0119] Example 1

[0120] Figure 3 The diagram shows the distribution of the indoor microphone array and sound source locations used in this example. Let the room length, width, and height be x, respectively. room y room h room The reference microphone position is (x ref ,y ref ,h ref The location of the sound source is (x) s ,y s ,h s ), h s =h ref The speed of sound in air is c = 340 m / s. A uniform linear array of four omnidirectional microphones (M1, M2, M3, M4) is used to receive the speech signal, with an element spacing D = 0.04 m. M1 is used as the reference channel, and the distance between the sound source and the reference microphone is r, with an azimuth angle of θ. A room reverberation model is established using the virtual source method. The signal sampling frequency is set to 8 kHz, the frame length to 10 ms, and the frame shift to half the frame length, resulting in 80 sampling points W per frame of speech signal. Speakers' voices are randomly selected from the TIMIT speech library as the sound source signal, ensuring that the speech in the training, validation, and test sets is different, with proportions of 65%, 10%, and 25%, respectively. Assuming that the additive noise between channels is uncorrelated, nine types of noise are selected from the Freesound sound effects library and white noise from the NOISEX-92 noise library during training and validation, and these are added to the reverberant speech received by the microphone array. During the testing process, three noise signals—babble, destroyerops, and factory—were selected from the NOISEX-92 noise library and added to reverberant speech. Data was generated under various conditions, including different combinations of reverberation time (RT60), signal-to-noise ratio (SNR), noise type, room size, microphone array location, sound source location, and speaker's speech signal. The reverberation time (RT60) and room length x [the remaining parameters] were used during training, validation, and testing. room Room width y room Room height h room The distance r between the sound source and the reference microphone, and the azimuth angle θ, are randomly selected within the ranges of [0.2, 0.6]s, [6.5, 9.5]m, [3.5, 6.5]m, [2.5, 3.5]m, [1.5, 2]m, and [0, 180]°, respectively. The signal-to-noise ratio (SNR) during training and validation is randomly selected within the range of [-5, 5]dB. The SNR during testing is fixed at 0dB and 5dB. The distance between the reference microphone and the center of the room is randomly selected within the range of [0, 0.5]m. The sound source is at least 0.5m away from the room wall.

[0121] Table 1 compares the network performance of the proposed method (denoted as OBFNet) and the filtering summation network FasNet under non-causal feature extraction and filtering conditions, using TCN and DPRNN for beamformer weight estimation at a filter length L=51. Because the proposed method considers contextual information during feature extraction, the speech enhancement effect under non-causal feature extraction is more significant. Under the same filter order, the proposed method demonstrates superior performance compared to FasNet across all evaluation metrics. Compared to TCN, which has a fixed receptive field, DPRNN can fully utilize intra-frame and inter-frame information, resulting in better network training performance. This trend remains unchanged after adding degree-of-freedom extension, indicating that the proposed method is applicable to speech enhancement of various temporal networks.

[0122] Table 1

[0123]

[0124] Comparative Example 1

[0125] Similar to the data generation method in Example 1, Table 2 shows a comparison of the speech enhancement performance of the method of the present invention and FasNet under different filter lengths L, using DPRNN to train beamformer weights, in the case of non-causal feature extraction and filtering. Clearly, the method of the present invention can achieve the same speech enhancement effect as FasNet with a longer filter length even with a shorter filter length, while simultaneously reducing network algorithm latency.

[0126] Table 2

[0127]

[0128]

[0129] Comparative Example 2

[0130] In this comparative example, the frame length of the short-time Fourier transform in the frequency domain method is set to 64ms, and the frame shift is half the frame length. Similar to the data generation method described in Example 1, Table 3 shows a comparison of the speech enhancement performance of the method of the present invention with traditional beamforming methods TI-MVDR and BeamformIt, and existing DL-based frequency domain beamforming methods GRN-IRM-MVDR and BeamTasNet-sigMVDR, when the filter length L = 51, under non-causal feature extraction and filtering conditions. Compared to frequency domain methods, the method of the present invention, as a time domain method, can exhibit good noise reduction performance with shorter frame lengths and filter lengths, with the OBFNet-DPRNN method showing particularly significant advantages.

[0131] Table 3

[0132] method △PESQ △STOI (%) SI-SNR (dB) SDR (dB) Oracle TI-MVDR 0.45 14.91 7.41 8.53 BeamformIt 0.20 5.72 4.13 5.56 GRN-IRM-MVDR 0.37 9.33 7.38 8.66 BeamTasNet-sigMVDR 0.51 12.70 8.55 9.87 OBFNet-TCN 0.85 19.73 11.43 12.33 OBFNet-DPRNN 0.92 20.76 11.82 12.62

Claims

1. A low-latency speech enhancement method based on a microphone array, characterized in that, Includes the following steps: Step 1: Set a set of initial pole parameters, optimize the initial pole parameters using an artificial neural network to obtain real poles, and use these real poles as parameters for the multi-pole orthogonal basis function model of each channel of the microphone array; Step 2: Construct orthogonal basis function models for each channel of the microphone array using real poles, and calculate the responses of each order of filters; Step 3 involves framing and extracting temporal features from the microphone array received signal, and then using an improved temporal network to estimate the weights of the adaptive beamformer constructed from orthogonal basis function models; the specific process is as follows: If the microphone received signal is divided into frames with a frame length of W and a frame shift of J = W / 2, then the k-th frame signal of the m-th channel of the microphone array is represented as: u m (k)=[u m ((k-1)J+1),…,u m ((k-1)J+W)] T In the case of non-causal feature extraction and filtering, concatenating the current frame signal and its context information yields the following concatenated signal: v m (k)=[u m ((k-1)J-C+1),…,u m ((k-1)J+W+C)] T In the formula, C represents the length of the preceding or following information; Signal u is received via reference microphone ref (k) and v m (k) Extract the feature vector ξ of each channel of the microphone array m (k), and the eigenvector ξ m (k) is input into the improved temporal network SeqNet(·) to obtain the weights of the adaptive wave velocity shaper; The improved temporal network SeqNet(·) begins with a linear bottleneck layer of B output channels. Therefore, the weights of the frame-level adaptive beamformers corresponding to the state-space equations are estimated as follows: w m (k)=OutputLayer(SeqNet(ξ m (k))) In the formula, w m (k)=[w m,1 (k),…,w m,L (k)] T This is the composition vector of the adaptive wave velocity generator weights; The output layer (OutputLayer) of the improved temporal network SeqNet(·) consists of a one-dimensional convolutional neural network and an activation function, with the following specific settings: OutputLayer(p)=PReLU(n(Wp+q)) In the formula, This is the output of a time-series network. and , where are the weights and biases of the one-dimensional convolutional neural network, η is the scale factor that controls the output of the one-dimensional convolutional neural network, and PReLU(·) is the activation function of the parameter-corrected linear unit. Step 4: Calculate the system response of each channel of the beamforming network based on the filter response and beamformer weights, and filter and sum the received signal of the microphone array to obtain the enhanced frame-level speech signal. Perform an overlap and summation operation on the enhanced frame-level speech signal to obtain the enhanced complete speech signal.

2. The low-latency speech enhancement method based on a microphone array according to claim 1, characterized in that, Step 1 includes: For a microphone array with m channels, m = 1, 2, ..., M, an unbiased one-dimensional convolutional neural network with 1 input channel and L output channels is used to train the initial pole parameters. M represents the number of microphones, L represents the filter length, and the kernel size and stride of the one-dimensional convolutional neural network are both set to 1. Data with a constant value of 1 is used as input, and the output is the learnable weights of this one-dimensional convolutional neural network, expressed as: In the formula, Conv1d represents a one-dimensional convolutional neural network; l represents the filter order number; The extreme points obtained in each training session Perform boundary processing, and then process the pole a. ml As poles of the orthogonal basis function model; where the boundary treatment is expressed as: In the formula, HardTanh(·) is the activation function, expressed as: In the formula, a max and a min These represent the maximum and minimum values ​​of the linear range of the HardTanh(·) function, respectively.

3. The low-latency speech enhancement method based on a microphone array according to claim 1, characterized in that, Step 2 includes: The state-space equation of the orthogonal basis function model of the m-th channel of the microphone array is: x m (t)=A m x m (t-1)+b m u m (t) In the formula, Let be the state vector at the t-th discrete time point. The state matrix, For the input-state vector, u m (t) represents the signal received by the microphone, w m This refers to the state-output vector, i.e., the weights of the orthogonal basis function model. For the output of the orthogonal basis function model, [·] T Represents the transpose of a matrix or vector; Calculate the state matrix A based on the state-space equation of the orthogonal basis function model. m The (j,l) elements and the input-state vector b m The l-th element is: in, This is the orthogonal state-space realization matrix of the l-th first-order all-pass filter that forms the orthogonal basis function model of the m-th channel, and the elements in the matrix are related to the poles a. ml The relationship is: d ml =-a ml ; Set the initial state vector x m If (0) = 0, then the state-space equation of the orthogonal basis function model is: In the formula, i represents the discrete-time index. Represents the state matrix A m The i-1th power; Therefore, the length is L. h The response of the L-order filter is: in, Let L represent the response vector of the L-th filter at time point i, where i = 1, 2, ..., L. h L h ≥L.

4. The low-latency speech enhancement method based on a microphone array according to claim 1, characterized in that, In step 4, the system response of each channel of the beamforming network is as follows: When L h When the value is greater than L, pad with L before convolutional concatenation. h -L zeros are added to ensure that the length of the convolutional signal is equal to the expected signal length. The concatenated information after zero padding is then: A one-dimensional convolutional neural network is used to implement beamforming network filtering of the received signals of each channel, and the output signals of all channels are summed to obtain the enhanced k-th frame speech signal: In the formula, * represents the convolution operation of a one-dimensional convolutional neural network, and the relationship between the filter length and the context signal length is L = 2C + 1.

Citation Information

Patent Citations

  • Microphone array speech enhancement method and device, electronic equipment and storage medium

    CN113889137A

  • Design method of frequency-invariant broadband beam former with adjustable main lobe direction

    CN118214979A