Semi-blind source separation method for nonlinear acoustic echo cancellation

By using a semi-blind source separation method based on short-time Fourier transform of convolutional transfer function and independent vector analysis or independent low-rank matrix analysis of auxiliary function, the problems of excessive time delay and improper step size adjustment in the prior art are solved, and better nonlinear acoustic echo cancellation effect is achieved in high reverberation environment.

CN115294996BActive Publication Date: 2025-12-05NANJING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210684642.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2025-12-05
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

Existing nonlinear acoustic echo cancellation methods suffer from excessive time delay and improper step size adjustment in real-time applications, resulting in poor performance in high reverberation environments. In particular, methods based on semi-blind source separation are difficult to apply effectively in nonlinear acoustic echo cancellation.

Method used

A semi-blind source separation method based on short-time Fourier transform of convolutional transfer function and independent vector analysis or independent low-rank matrix analysis of auxiliary function is adopted. By optimizing the separation matrix, the near-end time-frequency domain signal is separated, reducing time delay and improving nonlinear echo cancellation performance.

Benefits of technology

It achieves better nonlinear echo cancellation performance under real-time conditions, especially in high reverberation environments, showing superior performance compared to existing methods, thus improving echo cancellation and near-end speech quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294996B_ABST
    Figure CN115294996B_ABST
Patent Text Reader

Abstract

The application discloses a semi-blind source separation method for nonlinear acoustic echo cancellation. The method comprises the following steps: (1) obtaining a microphone signal containing nonlinear echo to be processed; (2) performing basis function expansion on a nonlinear mapping input signal, and obtaining a time-frequency domain observation model by using a short-time Fourier transform based on a convolution transfer function approximation, so as to obtain a semi-blind source separation model for nonlinear acoustic echo cancellation; (3) according to the semi-blind source separation model, realizing semi-blind source separation of the signal based on an auxiliary function independent vector analysis method or an independent low-rank matrix analysis method, optimizing a separation matrix, and separating out a near-end time-frequency domain signal; and (4) obtaining a time-domain near-end signal by using an inverse short-time Fourier transform. The method has effective nonlinear echo cancellation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of signal processing, specifically relating to a semi-blind source separation method for nonlinear acoustic echo cancellation. Background Technology

[0002] Acoustic echo cancellation (AEC) has been a very important research area for decades. Commonly used linear AEC methods can remove linear echoes from systems, but their performance becomes insufficient when nonlinear distortion in the system is non-negligible. Therefore, nonlinear AEC (NAEC) has received increasing attention, requiring not only the estimation of the echo path but also the modeling of the loudspeaker's nonlinearity.

[0003] The recently proposed NAEC method based on semi-blind source separation (SBSS) exhibits better nonlinear echo cancellation performance compared to traditional NAEC methods based on adaptive filtering. However, the proposed SBSS method suffers from two weaknesses in real-time applications. First, the SBSS method operates in the Short-Time Fourier Transform (STFT) domain based on the Multiplicative Transfer Function (MTF) approximation. The MTF approximation relies on the assumption that the STFT window length is sufficiently large compared to the length of the system impulse response. A long analysis window length leads to long time delays, thus making MTF-based SBSS unsuitable for real-time NAEC, especially in high-reverberation environments. Second, the separation matrix of the SBSS method is updated using the Natural Gradient algorithm, commonly used in Independent Vector Analysis (IVA). However, the Natural Gradient algorithm struggles to determine a suitable step size to effectively balance convergence speed and stability. In the field of Blind Source Separation (BSS), Auxiliary-function-based IVA (AuxIVA) and Independent Low-Rank Matrix Analysis (ILRMA) have been shown to outperform Natural Gradient IVA (NGIVA) because their iterative update algorithms do not require a step size and have more flexible source models. However, due to the special structure of the separation matrix, their optimization schemes cannot be directly applied to BSS. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a semi-blind source separation method for nonlinear acoustic echo cancellation based on auxiliary function independent vector analysis or independent low-rank matrix analysis.

[0005] The technical solution adopted in this invention is as follows:

[0006] A semi-blind source separation method for nonlinear acoustic echo cancellation includes the following steps:

[0007] Step 1: Acquire the microphone signal containing nonlinear echo to be processed;

[0008] Step 2: Expand the basis function of the nonlinear mapped input signal and obtain the time-frequency domain observation model using the short-time Fourier transform based on the convolution transfer function approximation, thereby obtaining the semi-blind source separation model for nonlinear acoustic echo cancellation.

[0009] Step 3: Based on the semi-blind source separation model, perform semi-blind source separation of the signal using the auxiliary function independent vector analysis method or the independent low-rank matrix analysis method, optimize the separation matrix, and separate the near-end time-frequency domain signal.

[0010] Step 4: Obtain the near-end signal in the time domain through short-time inverse Fourier transform.

[0011] Compared with existing technologies, the method of the present invention utilizes the convolution transfer function approximation to make semi-blind source separation applicable to real-time echo cancellation, and enables auxiliary function independent vector analysis and independent low-rank matrix analysis to be used in semi-blind source separation methods, thereby obtaining better nonlinear echo cancellation performance. Attached Figure Description

[0012] Figure 1 This invention provides a CTF-based SBSS model for NAEC.

[0013] Figure 2 The ERLE performance of the existing SBSS-NGIVA method based on MTF and CTF respectively and the method of the present invention are compared under various reverberation conditions in the case of single-ended calls.

[0014] Figure 3 The tERLE performance of the existing SBSS-NGIVA method based on MTF and CTF respectively in the case of two-way communication is compared with that of the present invention under various reverberation conditions.

[0015] Figure 4 The existing SBSS-NGIVA method based on CTF in single-ended call scenarios differs from the method of this invention in T 60 ERLE performance at 0.3s.

[0016] Figure 5The existing SBSS-NGIVA method based on CTF in the case of two-way calls differs from the method of this invention in T 60 tERLE performance at 0.3s.

[0017] Figure 6 The ERLE performance of the existing SBSS-NGIVA method based on CTF in single-ended call scenarios and the method of this invention when using recorded signals is compared.

[0018] Figure 7 The tERLE performance of the existing SBSS-NGIVA method based on CTF in a two-way call scenario is compared with that of the method of the present invention when using recorded signals. Detailed Implementation

[0019] The semi-blind source separation method for nonlinear acoustic echo cancellation of this invention mainly includes the following steps:

[0020] 1. Signal Acquisition

[0021] Acquire the microphone signal y(t) containing nonlinear echo to be processed, such as Figure 1 As shown, y(t) can be expressed as:

[0022] y(t)=d(t)+s(t)=h(t)*f(x(t))+s(t) (1)

[0023] Where t is time, d(t) is the echo signal, s(t) is the near-end signal, h(t) is the echo path, x(t) is the far-end input signal, f(·) represents the nonlinear mapping function, and f(x(t)) is the nonlinear mapping input signal.

[0024] 2. Construct an SBSS model based on Convolutive Transfer Function (CTF) for NAEC.

[0025] 1) Perform basis function expansion on the nonlinear mapped input signal:

[0026]

[0027] Where, φ i (·)=(·) 2i-1 Let a be the i-th order basis function. i Here, P represents the expansion coefficients, and P represents the expansion order. Substituting equation (2) into equation (1), we obtain the time-domain observation model as follows:

[0028]

[0029] 2) CTF-based STFT

[0030] The CTF approximation can represent long-impulse responses in the STFT domain using short-time frames, reducing latency compared to the MTF approximation and making it more suitable for real-time NAEC. Using the CTF-based STFT, the time-frequency domain observation model can be obtained as follows:

[0031]

[0032] Where k is the frequency index, n is the frame index, L is the number of short-time frames, Y(k,n), X φ,i (k,n) and S(k,n) are y(t) and φ respectively. i The time-frequency domain representations of (x(t)) and s(t), H l (k,n) is the acoustic transfer function, H i,l (k,n)=a i H l (k,n) is a hybrid nonlinear acoustic transfer function.

[0033] 3) SBSS model

[0034] Define an L×1 vector h i (k,n)=[H i,0 (k,n),…,H i,L-1 (k,n)] T and x i (k,n)=[X φ,i (k,n),…,X φ,i (k,n-L+1)] T ,in(·) T If we denote the transpose, then equation (4) can be written as:

[0035]

[0036] Redefine the (PL+1)×1 mixed output vector and mixed source vectors Then equation (5) can be written in the form of a vector matrix:

[0037] y(k,n)=H(k,n)s(k,n) (6)

[0038]

[0039] Where H(k,n) is a (PL+1)×(PL+1) mixing matrix. A PL×1 blend vector, 0 PL×1 Let I be the zero vector of PL×1. PL It is a PL×PL identity matrix.

[0040] The SBSS method can be used to obtain the estimated value E(k,n) of the near-end time-frequency domain signal S(k,n):

[0041] e(k,n)=W(k,n)y(k,n) (8)

[0042]

[0043] in, Let W(k,n) be the estimated vector of (PL+1)×1, W(k,n) be the separation matrix of (PL+1)×(PL+1), and w(k,n) be the separation vector of PL×1.

[0044] 3. SBSS Method

[0045] 1) SBSS method based on AuxIVA

[0046] In the SBSS model, due to the constraint structure of the separation matrix W(k,n), only the first row of W(k,n) is not fixed, while the other rows are fixed. Therefore, only the first row needs to be iterated, while keeping the first element as 1. Simultaneously, online AuxIVA is required for real-time NAEC. Therefore, the iterative formula for optimizing the separation matrix using the AuxIVA-based SBSS method is as follows:

[0047]

[0048] V1(k,n)=αV1(k,n-1)+(1-α)Φ(r1(n))y(k,n)y H (k,n) (11)

[0049]

[0050]

[0051] Where K is the total number of frequency points, (·) H This represents the conjugate transpose, where α is the forgetting factor. The first row of W(k,n), w 1,1 (k,n) is the first element of w1(k,n), r1(n) is an auxiliary variable, and Φ(r1(n)) = r1 β-2 (n) is the weighting function, β is the shape parameter, V1(k,n) is the weighted covariance matrix, and e1 represents the one-hot vector with the first element being 1.

[0052] 2) SBSS method based on ILRMA

[0053] Similar to the AuxIVA-based SBSS method, the online ILRMA method iteratively updates only the first row of W(k,n). Therefore, the iterative formula for optimizing the separation matrix using the ILRMA-based SBSS method is:

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060] Where b is the basis index, B is the number of basis elements, e1(k,n) is the first element of the estimated vector e(k,n), t1(k,b) is the basis of the source, v1(b,n) is the source activation, and r1(k,n) is the estimated source variance, which is calculated once after each update of t1(k,b) and v1(b,n).

[0061] 3) Reconstructing the near-end signal in the time domain

[0062] Finally, the near-end time-frequency domain signal separated by the SBSS method is used to obtain the near-end time domain signal through short-time Fourier inverse transform.

[0063] Example

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0065] 1. Parameter settings

[0066] The short-time Fourier transform uses a Hanning window with a window length of 1024, a frame shift of 256, and a sampling rate of 16 kHz. The nonlinear expansion order P and the number of short-time frames L are both 3, the forgetting factor α is 0.99, the shape parameter β is 0.4, and the number of bases B is 10.

[0067] 2. Simulation Examples

[0068] 1) Nonlinear mapping

[0069] A hard clipping model is used to generate the nonlinear mapped input signal:

[0070]

[0071] The clipping threshold is set to x. max =0.2max|x(t)|.

[0072] 2) Test Samples

[0073] The echo path is simulated using the room impulse response generated by the mirror method, with a reverberation time T. 60 The time interval is set from 0.2s to 0.8s, with a step size of 0.1s. Gaussian white noise is used as the background noise, with a signal-to-noise ratio of 60dB. Both single-ended and two-ended call scenarios are considered, with a 10s long male voice signal used as the far-end input signal in both cases. For two-ended calls, a 10s long female voice signal is used as the near-end signal, and the signal-to-echo ratio (SER) is set to 0dB.

[0074] 3) Evaluation Indicators

[0075] For one-way calls, the algorithm performance is evaluated using Echo Return Loss Enhancement (ERLE), defined as 10log 10 {E[y 2 (t)] / E[e 2 (t)]}. For two-way calls, true echo loss enhancement (true ERLE, tERLE) is used, defined as 10log 10 {E[d 2 (t)] / E[(e(t)-s(t)) 2 Near-end speech quality was evaluated using both Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI).

[0076] 4) Simulation results

[0077] To demonstrate the performance of the method of this invention, this embodiment will be compared with the existing SBSS method based on natural gradient IVA (G.Cheng, L.Liao, H.Chen, and J.Lu, “Semi-blind source separation for nonlinearacoustic echo cancellation,” IEEE Signal Process. Lett., vol.28, pp.474–478, Feb.2021.). This method uses both the original MTF model (L=1) and the CTF model of this invention. Figure 2 and Figure 3The ERLE and tERLE performances of the CTF-based SBSS-AuxIVA and SBSS-ILRMA methods of this invention, and the existing SBSS-NGIVA methods based on MTF and CTF, respectively, are presented under various reverberation conditions in single-ended and two-ended call scenarios. It can be seen that the CTF-based SBSS-NGIVA consistently outperforms the MTF-based SBSS-NGIVA, especially under high reverberation conditions. Furthermore, the CTF-based SBSS-AuxIVA and SBSS-ILRMA methods of this invention exhibit better performance than the CTF-based SBSS-NGIVA. Figure 4 and Figure 5 Three CTF-based SBSS methods are presented in T 60 The ERLE and tERLE performance at 0.3s, and the PESQ and STOI results during two-way communication are shown in Table 1. This demonstrates the advantages of the SBSS-AuxIVA and SBSS-ILRMA methods of this invention in terms of echo cancellation and near-end speech quality.

[0078] Table 1. PESQ and STOI results during two-way communication.

[0079]

[0080] 3. Experimental Examples

[0081] The performance of the method of the present invention was evaluated using a recorded signal. A low-cost, small loudspeaker was used to achieve nonlinear echo, and a 10-second female voice signal with a reverberation time of approximately 0.5 seconds was recorded using a microphone in an office setting. In a two-way call scenario, a 10-second male voice signal was used as the near-end signal, with a SER of 0 dB. Figure 6 and Figure 7 The ERLE and tERLE performances of three CTF-based SBSS methods are presented for single-ended and two-ended calls, respectively. The PESQ and STOI results for two-ended calls are shown in Table 2. The effectiveness of the proposed method can also be seen from the recorded signal results.

[0082] Table 2 shows the PESQ and STOI results of the recorded signals during two-way communication.

[0083]

Claims

1. A semi-blind source separation method for nonlinear acoustic echo cancellation, characterized in that, The method comprises the following steps: Step 1, obtaining a microphone signal y(t) to be processed containing nonlinear echo; Step 2, performing basis function expansion on a nonlinear mapping input signal, and obtaining a time-frequency domain observation model by using a short-time Fourier transform based on a convolution transfer function approximation, to obtain a semi-blind source separation model for nonlinear acoustic echo cancellation; the time-frequency domain observation model is: where k is the frequency index, n is the frame index, L is the number of short-time frames, Y(k, n), X φ,i (k, n) and S(k, n) are the time-frequency domain representations of the microphone signal y(t), the i-th basis function φ i (x(t)) and the near-end signal s(t), respectively, H l (k, n) is the acoustic transfer function, H i,l (k, n) = a i H l (k, n) is the mixed nonlinear acoustic transfer function, a i is the expansion coefficient, and P is the expansion order; Step 3, performing semi-blind source separation of the signal based on an auxiliary function independent vector analysis method or an independent low-rank matrix analysis method according to the semi-blind source separation model, optimizing a separation matrix, and separating out a time-frequency domain near-end signal; Step 4, obtaining a time domain near-end signal by inverse short-time Fourier transform.

2. The semi-blind source separation method for nonlinear acoustic echo cancellation according to claim 1, characterized in that, In the step 1, the microphone signal to be processed containing nonlinear echo is represented as: y(t) = d(t) + s(t) = h(t) * f(x(t)) + s(t) Wherein, t is time, d(t) is an echo signal, s(t) is a near-end signal, h(t) is an echo path, x(t) is a far-end input signal, f(·) represents a nonlinear mapping function, and f(x(t)) is a nonlinear mapping input signal.

3. The semi-blind source separation method for nonlinear acoustic echo cancellation of claim 1, wherein, In the step 2, the basis function expansion of the nonlinear mapping input signal is: where φ i (·) = (·) 2i-1 are the basis functions of order i, a i are the corresponding expansion coefficients, and P is the expansion order.

4. The semi-blind source separation method for nonlinear acoustic echo cancellation of claim 1, wherein, In the step 3, the iteration formula for performing semi-blind source separation of the signal based on the auxiliary function independent vector analysis method is: V1(k, n) = aV1(k, n - 1) + (1 - a)Φ(r1(n))y(k, n)y H (k, n) where k is the frequency index, n is the frame index, K is the total number of frequency points, (·) H denotes the conjugate transpose, a is the forgetting factor, y(k, n) is the mixed output vector, W(k, n) is the separation matrix, w H1 (k, n) is the first row of W(k, n), w 1,1 (k, n) is the first element of w1(k, n), r1(n) is an auxiliary variable, is a weighting function, β is a shape parameter, V1(k, n) is a weighted covariance matrix, e1 denotes a one-hot vector with the first element being 1.

5. The semi-blind source separation method for nonlinear acoustic echo cancellation of claim 1, wherein, In the step 3, the iteration formula for performing semi-blind source separation of the signal based on the independent low-rank matrix analysis method is: Wherein, b is a basis sequence number, B is the number of bases, e1(k,n) is the first element of an estimated vector e(k,n), t1(k,b) is a basis of a source, v1(b,n) is a source activation, r1(k,n) is an estimated source variance, and is calculated once after each update of t1(k,b) and v1(b,n).

Citation Information

Patent Citations

  • Non-linear acoustic echo cancellation method based on semi-blind source separation

    CN112927706A

  • Multi-channel non-negative matrix factorization method and system based on frequency domain convolution transfer function

    CN114220453A