An echo cancellation device and an echo cancellation method based on frequency-domain block IPNLMS

Through the microphone array and frequency domain blocking IPNLMS algorithm combined with beamforming technology, the problem of acoustic echo interference in far-field voice interaction is solved, efficient echo cancellation and reduction in computational complexity, and improved voice signal quality.

CN115938381BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211607768.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-07-11
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

In far-field voice interaction scenarios, ambient noise, reverb and multi-source crosstalk lead to a decrease in the signal-to-noise ratio of the voice signal. The traditional echo cancellation algorithm converges slowly and has a large amount of calculation in the dual-talk state, making it difficult to effectively reduce acoustic echo interference in limited resource equipment.

Method used

Voice preprocessing is performed using microphone arrays, combined with frequency domain blocking IPNLMS algorithm and beamforming technology, and block-by-block update and time-varying step factor are reduced by adaptive filters.

Benefits of technology

While ensuring the echo cancellation performance, it significantly reduces the computational complexity, improves convergence speed, adapts to environmental changes, and improves voice signal quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938381B_ABST
    Figure CN115938381B_ABST
Patent Text Reader

Abstract

The present invention discloses an echo cancellation device based on frequency-domain block IPNLMS, which includes a microphone array, a speech preprocessing module, and an echo cancellation module; based on the microphone array architecture, a beamforming method is used to preprocess the speech signal; the echo cancellation module adopts a block-by-block update method and executes an improved proportionate normalized least mean square algorithm in the frequency domain. The echo cancellation method of the present invention uses the fast Fourier transform to achieve time-domain linear convolution and linear correlation, solves the disadvantage of large computational complexity caused by point-by-point update in the original time-domain calculation, and at the same time introduces a time-varying step factor to have a superior algorithm convergence speed for the echo path with sparse characteristics. In the signal preprocessing stage of the present invention, the microphone array architecture is adopted as the speech acquisition module to achieve the spatial selectivity of the signal, effectively reducing the acoustic echo interference caused by the coupling between the microphone and the speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an echo cancellation device, and more particularly to an echo cancellation device and an echo cancellation method based on frequency-domain block IPNLMS. Background Art

[0002] Current speech processing technologies have reached a practical level for human-machine close-range interaction, and can collect high-quality speech signals in processing near-field speech signals. However, for long-distance interactions several meters away, especially in indoor far-field speech interaction scenarios, due to the influence of environmental noise, reverberation, and multi-source crosstalk, the signal-to-noise ratio of the speech signal drops severely, resulting in poor quality of the collected speech signal. Therefore, traditional hands-free phones or teleconferences generally require the speaker to be close to the speech terminal and have high requirements for the call environment of both parties to ensure high-quality voice calls.

[0003] The hands-free call process can be summarized into three states: the state where the proximal person is speaking, the state where the distal person is speaking, and the double-talk state. In the double-talk state, due to the coupling between the microphone and the speaker, the problem of acoustic echo interference will inevitably occur. This coupling causes the distal speech signal to be picked up by the proximal microphone when it is output by the proximal speaker, resulting in the distal speaker being able to hear their own voice. Therefore, compared with the state where only one party is speaking, the difficulty of acoustic echo cancellation is more manifested in the speech communication problem in the double-talk state.

[0004] The adaptive filter is one of the mainstream solutions for echo cancellation. The least mean square (LMS) algorithm proposed by Widrow and Hoff in the early stage has been widely used in acoustic fields such as speech recognition and echo cancellation because of its easy understanding and small computational complexity. However, this algorithm uses a fixed step-size factor, resulting in a very slow convergence rate. The later proposed normalized least mean square (NLMS) algorithm can select a variable step-size factor according to the input signal, and the convergence rate is greatly improved when the input signal is irrelevant. However, the speech signal is a complex time-varying signal with strong correlation between frames, which is mainly reflected in the co-articulation phenomenon of speech. In addition, the sparsity of the acoustic echo path is also one of the reasons for the performance degradation of the NLMS algorithm. The improved proportionate normalized least mean square (IPNLMS) algorithm makes the step size proportional to the estimated echo path by introducing a coefficient matrix, which improves the convergence characteristics of the algorithm to a certain extent, but the computational complexity increases.

[0005] In addition, compared with the early single-microphone pickup, the microphone array architecture can make full use of the spatial characteristics of speech signals to achieve spatial selectivity of signals. For a hands-free phone with fixed microphone and speaker positions, by using the beamforming technology of the microphone array, signal suppression can be performed on the direction of the speaker or the direction where other noises are generated, and the signal in the direction of the speaker can be enhanced, thereby effectively reducing the impact of acoustic echo interference in the double-talk state.

[0006] With the advent of the big data era, multimedia data has grown explosively, and the processing volume of audio and video is also extremely large. Nowadays, various small, portable, and low-cost intelligent voice terminal devices have increasingly entered people's lives. How to use efficient echo cancellation algorithms and sound source localization technologies in voice devices with limited processor computing power and memory space to handle the noise cancellation problems from the environment and the device, and ensure the quality of voice signals under both near-field and far-field conditions, is a good manifestation of the implementation of advanced voice processing technologies today. Summary of the Invention

[0007] Object of the Invention: The object of the present invention is to provide an echo cancellation device and an echo cancellation method based on frequency-domain block IPNLMS that can effectively reduce the acoustic echo interference caused by the coupling between the microphone and the speaker.

[0008] Technical Solution: The echo cancellation device of the present invention includes:

[0009] A microphone array, which is a voice preprocessing module and serves as the hardware architecture for beamforming in voice preprocessing. It uses multiple array elements to collect voices.

[0010] A voice preprocessing module for preprocessing the voice before echo cancellation. The voice preprocessing module adopts a beamforming method and is implemented by multiple phase delay groups and a multi-source selector. Each phase delay group consists of multiple phase delayers. The phase delayer is used to align the voice signals collected by multiple array elements in the microphone array. The multi-source selector selects the best source based on the power size of the transmitted source voice signal as the measurement criterion.

[0011] An echo cancellation module for further echo cancellation of the voice signal. The echo cancellation module includes a serial-parallel / parallel-serial converter, an adaptive filter, and a double-talk detector. The serial-parallel / parallel-serial converter is used to block and merge the signal. The adaptive filter uses the frequency-domain block IPNLMS algorithm to perform block processing on the preprocessed signal and updates the adaptive filter coefficients block by block. The double-talk detector detects whether there is double-end talking to control the working state of the filter.

[0012] Furthermore, the speech preprocessing module introduces delay-and-sum beamforming. By aligning the phases of the signals collected by each element of the microphone array according to the angle, further weighted summation and averaging are performed to form the output signal, and the azimuth angle of the sound source arrival is determined by the beam direction corresponding to the position of the maximum output power.

[0013] According to the azimuth angle of the sound source arrival, the multi-source selector in the speech preprocessing module extracts the signal corresponding to the angle as the input signal for the next echo cancellation module.

[0014] Furthermore, the ratio of the best source selected by the multi-source selector to the signal sources extracted by the beams at each angle is related to the sparsity of the estimated echo path; when the ratio of the L2-norms is close to 1, the step factor maintains the larger step at the previous discrete time.

[0015] Furthermore, the speech preprocessing module does not perform beamforming on the directly arriving echo signal from the speaker angle. When the number of elements of the microphone array is M, the multi-source selector weakens the original maximum echo signal power to 1 / M of the original.

[0016] Furthermore, the echo cancellation module determines whether there is double-talk according to the double-talk detector, and stops updating the filter coefficients when there is double-talk.

[0017] Furthermore, the echo cancellation module introduces a time-varying step size μ, which varies according to the change of the environmental sparsity.

[0018] An echo cancellation method for implementing echo cancellation of the above echo cancellation device, and the implementation process is as follows:

[0019] The microphone array uses multiple elements to collect speech.

[0020] The speech preprocessing module preprocesses the speech using the beamforming method.

[0021] The echo cancellation module performs block processing on the preprocessed signal using the frequency-domain block IPNLMS algorithm, and updates the adaptive filter coefficients block by block.

[0022] Among them, in the beamforming method, let the s-th phase delay group correspond to the beamforming at the α s angle, then the output of the s-th phase delay group is out s , then there is:

[0023]

[0024] Among them, M is the number of elements of the microphone array; x i is the speech signal collected by the i-th element, is the signal xi Time difference of arrival, 0 ≤ i ≤ M - 1;

[0025] The update of filter coefficients is carried out in the frequency domain. The detailed implementation process is as follows:

[0026] The input signal x(n) is divided into several data blocks of length L through a serial - parallel converter, and the filter coefficient update is performed only after accumulating L sampling points. The iterative formula for the adaptive filter coefficients is as follows:

[0027]

[0028] where μ is the global step size; δ > 0 is a constant; G(n) is a diagonal matrix used to add an independent scaling factor to each filter coefficient.

[0029] Compared with the prior art, the remarkable effects of the present invention are as follows:

[0030] 1. In the signal pre - processing stage, a microphone array architecture is adopted as the speech acquisition module to achieve spatial selectivity of the signal, effectively reducing the acoustic echo interference caused by the coupling between the microphone and the speaker; when the position of the near - end speaker changes or the environment changes, the sparsity of the echo path changes, and thus the step size of the adaptive filter coefficients is adjusted, with a faster convergence speed compared to the fixed step size;

[0031] 2. Adopting a block - by - block update method, the processed speech signal is input into the echo cancellation module, and the IPNLMS algorithm is executed in the frequency domain. The fast Fourier transform is used to realize time - domain linear convolution and linear correlation, solving the disadvantage of large computational complexity caused by point - by - point update in the original time - domain calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the structural block diagram of the echo cancellation device of the present invention;

[0033] Figure 2 is the internal structural block diagram of the speech pre - processing module;

[0034] Figure 3 is the schematic diagram of calculating the phase offset of adjacent two microphone array elements for the incident signal corresponding to the α angle;

[0035] Figure 4 is the schematic diagram of beam coherence for the target speech incident at the α angle;

[0036] Figure 5 is the schematic diagram of beam cancellation for the speaker echo incident at the β angle;

[0037] Figure 6 is the internal structural block diagram of the echo cancellation module. Detailed Implementation Manner

[0038] The present invention will be further described in detail below in conjunction with the accompanying drawings of the specification and the specific implementation manner.

[0039] The present invention provides an echo cancellation device that utilizes beamforming technology in cooperation with the IPNLMS algorithm with frequency-domain block division, which greatly reduces the computational complexity while ensuring the echo cancellation performance.

[0040] As Figure 1 shown, the echo cancellation device of the present invention mainly includes:

[0041] A microphone array that uses multiple array elements to collect the near-end speech signal as a speech acquisition module.

[0042] An analog-to-digital / digital-to-analog converter that converts continuous analog / digital signals into digital / analog signals.

[0043] A speech preprocessing module for speech processing before echo cancellation, including multiple phase delay groups and a multi-source selector, as Figure 2 shown. Each phase delay group consists of multiple phase delayers. The phase delayer is used to align the speech signals collected by multiple array elements in the microphone array. The multi-source selector selects the best source based on the power magnitude of the transmitted source speech signal.

[0044] An echo cancellation module, including a serial-to-parallel / parallel-to-serial converter, an adaptive filter, and a double-talk detector, as Figure 6 shown. The serial-to-parallel / parallel-to-serial converter is used to block and merge the signals; the adaptive filter adopts the frequency-domain block IPNLMS algorithm, and the working state of the filter is controlled by the double-talk detector to detect whether there is double-talk at both ends.

[0045] The implementation steps of the echo cancellation device of the present invention are as follows:

[0046] Step 1, determine the layout of the microphone array and the speaker, and design a speech preprocessing module including multiple phase delay groups. The internal framework of the speech preprocessing module is as Figure 2 shown. It includes the following steps:

[0047] Step 11, according to the distance between the array elements in the microphone array, perform phase delay calculation on the input signal with an incident angle of α, so as to enhance the speech signal incident at this angle. As Figure 3 shown, the distance between the first array element mic1 and the second array element mic2 is d0. For the speech signal with an incident angle of α (far-field model), the distance difference of the signal transmission between the two array elements is d1, and the arrival time difference is Δ t , so as to obtain the phase difference as Δ θ, By using the phase delay groups to perform phase offset on the received signals of each array element to achieve signal alignment, signal enhancement can be realized.

[0048] d1 = d0 × cosα (1)

[0049]

[0050] Among them, c is the speed of sound and f is the signal frequency.

[0051] Step 12, Repeat step 11, change the signal incident angle, calculate the new phase difference to design multiple phase delay groups for signal enhancement corresponding to multiple angles. Determine the position of the speaker in the device (assuming the incident angle of the echo signal output by the speaker is β), and do not design a phase delay group for this angle. Then, when other phase delay groups process the signal at this angle, the signal will be suppressed. The principle is as Figure 4 and Figure 5 shown. In this embodiment, a phase delay group is set every 30°, corresponding to the beamforming of this angle. At the same time, assume that β = 120° (that is, no phase delay group corresponding to the incident angle of 120° is set), then a total of 11 phase delay groups can be determined:

[0052] α = 0°, 30°, 60°, 90°, 150°, 180°, 210°, 240°, 270°, 300°, 330°.

[0053] Next is the working stage of the device.

[0054] Step 2, Voice acquisition and preprocessing. As Figure 1 shown, the microphone array collects the proximal signal, including the voice signal v(n) of the proximal speaker at discrete time n, the voice signal d(n) of the distal speaker output by the speaker, and the environmental noise signal r(n). Assume the equivalent filter h(n) of the entire indoor environment. Then, the voice signal x(n) collected by the microphone array is expressed as follows:

[0055] x(n) = h(n)*d(n) + v(n) + r(n) (4)

[0056] Step 21, Process the signals collected by multiple array elements in the microphone array by the voice preprocessing module designed in step 1. Let the s-th phase delay group correspond to the beamforming at the α s angle. Then, the output of the s-th phase delay group is out s .

[0057]

[0058] Among them, M is the number of array elements in the microphone array, x iis the voice signal collected by the i-th array element (0 ≤ i ≤ M - 1), Δ ti is the time difference of arrival of signal x i .

[0059] Step 22, based on the power magnitude of the signals, perform multi-source selection on the output signals of each phase delay group, and select the best source out max , and simultaneously obtain the azimuth angle α of the sound source arrival of the best source max .

[0060] out max (n) = max{out1(n), out2(n),..., out s (n),..., out 11 (n)} (6)

[0061] Here, out s represents the output signal of the s-th phase delay group, corresponding to the azimuth angle α of the sound source arrival s .

[0062] Step 3, calculate the time-varying step factor. The ratio of the L2-norm between the best source selected by the multi-source selector and the signal sources extracted by each angle beam is related to the sparsity of the estimated echo path. In an open environment, the reverberation time is significantly reduced, and the sparsity of the estimated echo path is large; while in an indoor environment, it is the opposite. When the azimuth angle α of the sound source arrival solved in Step 22 max changes, it can be considered that the echo path has changed, and thus the step factor is updated.

[0063]

[0064] where, ‖·‖2 represents the L2-norm operator. In addition, in the initial stage, take μ(0) = 0.4.

[0065] Step 4, echo path estimation. In the case where only the person at the far end is speaking, the task of the adaptive filter is to use the adaptive filter coefficients to estimate the echo path h(n), so that the estimated echo signal approximates the desired response signal d(n) as much as possible, and the error signal e(n) obtained by subtracting the two is used to update the filter coefficients, thereby realizing adaptive adjustment. Since the update of the filter coefficients only occurs when there is only the far-end signal, in this embodiment, the environmental noise signal r(n) is ignored, and the signal input at discrete time n is simplified to x(n) to replace the preprocessed voice input out max (n) in the real environment. Assume that the device uses an N-order filter. When the discrete time is n, the input signal x(n) and the filter coefficients can be respectively expressed as:

[0066] x(n) = [x(n), x(n - 1),..., x(n - N + 1)] T (8)

[0067]

[0068] Then the filter output is

[0069]

[0070] where "T" represents the transposed matrix.

[0071] Step 41, signal block processing. The input signal x(n) is divided into several data blocks of length L through a serial - to - parallel converter, and the filter coefficient update is performed only after accumulating L sampling points. The present invention uses an improved proportionate normalized least mean square algorithm (IPNLMS), which introduces a gain control matrix G(n) on the basis of the NLMS algorithm. After block processing, the adaptive filter coefficient iteration formula is as follows:

[0072]

[0073] where the gradient block:

[0074] where μ is the global step size of the algorithm. δ > 0, which is a constant. Its role is to prevent the denominator from approaching zero due to too small input signal, resulting in a decline in algorithm stability. G(n) is a diagonal matrix used to add an independent scale factor to each filter coefficient, so that the active components in the estimated echo path satisfying sparsity have a faster convergence rate, and it is expressed as:

[0075] G(n) = diag{g0(n), g1(n),..., g N-1 (n)} (13)

[0076] The diagonal elements are expressed as:

[0077]

[0078] where - 1 ≤ γ ≤ 1, which is an algorithm preference factor. It is easy to find that when γ = - 1, it is consistent with the adaptive process of the NLMS algorithm; when γ = 1, it is consistent with the PNLMS algorithm. Experiments show that γ is most suitable to take 0 or - 0.5. And ε is a small positive number, which is introduced to avoid division by zero in Equation (14) during initialization because all taps of the filter start from zero.

[0079] For unified representation, a new time index k is redefined in this embodiment here. Let n = kL, and at the same time let the data block length L = N. By eliminating the explicit dependence on L on both sides of the equation, Equation (12) can be rewritten as:

[0080]

[0081] Here, the new time index k is used to replace the discrete time moment n described previously to represent the filter coefficients is updated only once for each input block of data (each block has L input signal samples), and the update of G(k) is similar.

[0082] Step 42: Introduce the fast Fourier transform and use the overlap - save method to place the adaptive filter update in the frequency - domain calculation. According to the overlap - save method, the frequency - domain weight vector corresponds to a finite - length sequence, and the input signal corresponds to an infinite - length sequence. Each fast Fourier transform or inverse transform is implemented by a discrete Fourier transform (DFT) module. To calculate the 2N - point DFT, N zero components need to be appended to the time - domain weight vector with N components to form a 2N - order frequency - domain weight vector; for the input block signal x(k), the current data block and the previous data block are concatenated to form a 2N - order frequency - domain input signal. Therefore, the adaptive process derived from Equations (8) - (11) is transformed from the time - domain to the frequency - domain as follows:

[0083]

[0084] X(k) = diag{F[x(kN - N),...,x(kN - 1),x(kN),...,x(kN + N - 1)] T} (17)

[0085]

[0086] where H(k) is the frequency - domain representation of the filter coefficients, X(k) is the frequency - domain representation of the input signal, is the frequency - domain representation of the filter output, and E(k) is the frequency - domain representation of the error signal. F is defined as a 2N×2N - order DFT matrix, which is used to convert the time - domain quantities derived previously into frequency - domain quantities. Its inverse transform (IDFT) matrix is defined as F -1 .

[0087] According to the overlap - save method, the linear convolution of N output samples can be calculated as the last N components of , and the expression is as follows:

[0088]

[0089] For the filter coefficient update formula, before converting the linearly related calculations to the frequency domain for processing, it is still necessary to solve for the term value in the time domain. Let Then the block gradient in Equation (12) can be expressed as:

[0090]

[0091] Equation (12) can be abbreviated as:

[0092]

[0093] Then the solution of is linearly related to e(k), and the first N components of can be calculated. The expression is:

[0094]

[0095] where H represents the complex conjugate transpose.

Claims

1. An echo cancellation device based on frequency-domain block IPNLMS, characterized in that Including: A microphone array, which is a voice preprocessing module and serves as the hardware architecture for beamforming in voice preprocessing. It uses multiple array elements to collect voices. A voice preprocessing module for preprocessing voices before echo cancellation. The voice preprocessing module adopts a beamforming method and is implemented by multiple phase delay groups and a multi-source selector. Each phase delay group consists of multiple phase delayers. The phase delayer is used to align the voice signals collected by multiple array elements in the microphone array. The multi-source selector selects the best source based on the power magnitude of the transmitted source voice signal. An echo cancellation module for further echo cancellation of voice signals. The echo cancellation module includes a serial-parallel / parallel-serial converter, an adaptive filter, and a double-talk detector. The serial-parallel / parallel-serial converter is used to block and merge signals. The adaptive filter uses a frequency-domain block IPNLMS algorithm to perform block processing on the preprocessed signals and updates the adaptive filter coefficients block by block. The double-talk detector detects whether there is double-talk to control the working state of the filter. The echo cancellation module determines whether there is double-talk according to the double-talk detector. When there is double-talk, it stops updating the filter coefficients. And a time-varying step size μ is introduced, and the time-varying step size μ changes according to the change of the environmental sparsity.

2. The echo cancellation device based on frequency-domain block IPNLMS according to claim 1, wherein The voice preprocessing module introduces delay-and-sum beamforming. By phase-aligning the signals collected by each array element of the microphone array according to the angle, further weighted summation and averaging are performed to form an output signal, and the arrival azimuth angle of the sound source is determined by the beam direction corresponding to the position of the maximum output power. The voice preprocessing module extracts the signal corresponding to the angle as the input signal of the next echo cancellation module by the multi-source selector therein according to the arrival azimuth angle of the sound source.

3. The echo cancellation device based on frequency-domain block IPNLMS according to claim 2, characterized in that, The ratio of the L2-norm of the best source selected by the multi-source selector to the signal sources extracted by the beams of each angle is related to the sparsity of the estimated echo path. When the L2-norm ratio is close to 1, the step factor maintains the larger step size at the previous discrete moment.

4. The echo cancellation device based on frequency-domain block IPNLMS according to claim 2, characterized in that, The voice preprocessing module does not perform beamforming on the directly arriving echo signals from the speaker angle. When the number of array elements of the microphone array is M, the multi-source selector weakens the original maximum echo signal power to 1 / M of the original.

Citation Information

Patent Citations

  • Echo cancellation method and device, and call equipment

    CN106791244A

  • Variable-step-size hearing aid adaptive echo cancellation device and echo cancellation method

    CN111916099A