Acoustic echo cancellation using control parameters
By using the incremental maximum likelihood algorithm and parallel filter technology to dynamically adjust the confidence parameter, the problem of fast tracking and low residual echo in acoustic echo cancellation in multi-speaker and multi-microphone environments is solved, achieving robustness and low complexity for strong near-end signals.
Patent Information
- Application Number
- CN202210929674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-04
- Filing Date
- 2022-08-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-08-03
AI Technical Summary
In dynamic physical environments with multiple speakers and microphones, existing acoustic echo cancellation techniques struggle to achieve a balance between fast tracking, low residual echo, robustness to intermittent near-end signals, and low complexity.
The incremental maximum likelihood (IML) algorithm is adopted to dynamically update the filter coefficients of the echo cancellation system by adaptively adjusting the confidence parameter and combining speech activity detection and parallel filters, thereby achieving fast convergence and low residual echo.
Under different environmental conditions, it achieves rapid tracking of residual echo reduction, maintains low residual echo and robustness to strong near-end signals, while keeping computational complexity low.
Smart Images

Figure CN115706757B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The example embodiments herein generally relate to acoustic echo cancellation (AEC), and more specifically, to processes and apparatuses for performing AEC that can use maximum likelihood (ML) techniques. BACKGROUND
[0002] In a two-way audio system, there is typically a "far-end" and a "near-end." Consider a person in a room talking to a colleague in a different location via a video conference. The room is considered the "near-end" (relative to the person), and the location of the colleague is considered the "far-end."
[0003] Echo cancellation is required in any two-way audio system (e.g., a hands-free phone or conference room) where the loudspeaker and microphone are not physically isolated at the near-end, to prevent the far-end signal produced by the loudspeaker from feeding back to the far-end via the microphone. Such systems are now widely used, but new use cases involving spatial audio and immersive experiences make the technical problem more challenging.
[0004] Desirable properties of an audio echo cancellation system include one or more of the following:
[0005] 1) the ability to track rapidly changing physical environments even when the far-end signal is highly correlated;
[0006] 2) very low residual echo after convergence;
[0007] 3) robustness to the presence of intermittent strong near-end signals; and
[0008] 4) acceptable complexity (e.g., the length of the cancellation filter is linear). SUMMARY
[0009] This section is intended to include examples rather than limitations.
[0010] In one example embodiment, a method for echo cancellation for two-way audio communication is disclosed, the method including receiving, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal. One or more loudspeakers reproduce the far-end signal. The method includes operating the adaptive echo cancellation system with at least one filter at least in part to update an estimate of coefficients of a sound channel from the one or more loudspeakers to the one or more microphones. The method further includes determining at least one control parameter that affects operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values. The determination of the at least one control parameter is based on an accuracy of the estimate of the coefficients of the estimate of the sound channel and a characteristic of the near-end signal. The method includes controlling the at least one filter by the adaptive echo cancellation system at different times with different values of the at least one control parameter.
[0011] Another exemplary embodiment includes a computer program comprising code for performing the method of the preceding paragraph when the computer program is run on a processor. According to this paragraph, the computer program is a computer program product comprising a computer-readable medium bearing computer program code embodied therein for use with a computer. Another example is a computer program according to this paragraph, wherein the program is directly loadable into the internal memory of the computer.
[0012] An exemplary apparatus includes one or more processors and one or more memories including computer program code. The one or more memories and the computer program code are configured to, with the one or more processors, cause the apparatus to receive, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal, wherein the one or more loudspeakers reproduce the far-end signal; operate the adaptive echo cancellation system at least in part with at least one filter to update an estimate of a coefficient of a sound channel from the one or more loudspeakers to the one or more microphones; determine at least one control parameter affecting operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein determining the at least one control parameter is based on an accuracy of the estimate of the coefficient of the sound channel and a characteristic of the near-end signal; and control the at least one filter by the adaptive echo cancellation system at different times with different values of the at least one control parameter.
[0013] An exemplary computer program product includes a computer-readable storage medium bearing computer program code embodied therein for use with a computer. The computer program code includes code for receiving, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal, wherein the one or more loudspeakers reproduce the far-end signal; code for operating the adaptive echo cancellation system at least in part with at least one filter to update an estimate of a coefficient of a sound channel from the one or more loudspeakers to the one or more microphones; code for determining at least one control parameter affecting operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein determining the at least one control parameter is based on an accuracy of the estimate of the coefficient of the sound channel and a characteristic of the near-end signal; and code for controlling the at least one filter by the adaptive echo cancellation system at different times with different values of the at least one control parameter.
[0014] In another exemplary embodiment, an apparatus includes components for performing the following operations: receiving, at an adaptive echo cancellation system, an audio signal at one or more microphones, at least partially based on a near-end signal and a reproduced far-end signal, wherein one or more speakers reproduce the far-end signal; operating the adaptive echo cancellation system at least partially with at least one filter to update estimates of coefficients for channels from the one or more speakers to the one or more microphones; determining at least one control parameter affecting the operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one of a series of values, wherein determining the at least one control parameter is based on the accuracy of the estimates of the coefficients for the estimated channels and the characteristics of the near-end signal; and controlling at least one filter by the adaptive echo cancellation system at different times using different values of the at least one control parameter. Attached Figure Description
[0015] In the attached diagram:
[0016] Figure 1 This is a logic flowchart of an exemplary typical audio system with echo cancellation;
[0017] Figure 1A This is a block diagram of a communication device suitable for implementing echo cancellation, according to an exemplary embodiment;
[0018] Figure 2 This is a logic flowchart of the first embodiment, referred to as Embodiment 1;
[0019] Figure 3 This is a block diagram of an echo cancellation module, referred to as Embodiment 2, according to an exemplary embodiment;
[0020] Distributed in Figure 4A and 4B Figure 4 above is a logic flowchart of the general update process for Embodiment 2;
[0021] Figure 5 This is a logic flowchart of the periodic update rule in Example 2;
[0022] Figure 6 The illustration shows the evolution of signal power in a simulated SISO echo cancellation scenario using echo cancellation with two parallel filters running the IML algorithm in an exemplary embodiment.
[0023] Figure 7 The illustration shows echo cancellation using two parallel filters running the IML algorithm in an exemplary embodiment. Figure 6 Normalized misalignment in a simulated SISO echo cancellation scenario (20log) 10 ||w t -w *|| -20 log 10 || w * || evolution; and
[0024] Figure 8 is a logic flow diagram for acoustic echo cancellation using control parameters, and illustrates the operation of an example one or more methods, a result of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnected machine components for performing functions in accordance with the exemplary embodiments. DETAILED DESCRIPTION
[0025] Abbreviations that can be found in the specification and / or drawings are defined at the end of the DETAILED DESCRIPTION section below.
[0026] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations. The implementations described in this detailed description are intended to be merely illustrative of examples of implementations provided in connection with this disclosure, and are not intended to be limiting of the scope of the disclosure to the implementations described. For example, while the present disclosure is described in terms of specific implementations, it is contemplated that the applications described with respect to the implementations will have general applicability with regard to other implementations. Many modifications and variations of the described implementations are possible and will be apparent to those of ordinary skill in the art, in light of the above teachings. It is, therefore, to be understood that changes can be made to the implementations described and that aspects of the present disclosure can be arranged other than as described herein without departing from the spirit or scope of the underlying principles of the present disclosure. Accordingly, the application is intended to embrace all such alterations, modifications, and variations which fall within the scope of this disclosure, including full use of the explanatory material described herein.
[0027] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, elements and / or components but do not preclude the presence or addition of one or more other features, elements, components and / or groups thereof.
[0028] For system applications such as 5G immersive voice, it is desirable to utilize multiple loudspeakers and multiple microphones to provide a more realistic audio experience. For example, intelligibility can be enhanced by making different remote voices appear to come from different directions.
[0029] The use of multiple loudspeakers and microphones in large dynamic physical environments makes the acoustic echo cancellation problem more challenging for the following reasons:
[0030] 1) Multiple loudspeakers increase the correlation of the far-end signals, slowing down the convergence speed;
[0031] 2) Multiple microphones increase the computational complexity;
[0032] 3) Large physical environments increase the required length of the cancellation filter; and / or
[0033] 4) Dynamic physical environments increase the required tracking speed of the system.
[0034] To enable immersive voice applications, it would be useful to have an echo cancellation method that can simultaneously achieve fast convergence, low residual echo, robustness to near-end signal, and low complexity.
[0035] For the general problem of acoustic echo cancellation, there are many algorithms. Three key algorithms for adjusting the coefficients of the echo cancellation filter include the least mean square method (LMS), the recursive least squares method (RLS), and the affine projection algorithm (APA). While all of these are useful, they have the following limitations, which the technology presented herein seeks to address.
[0036] 1) LMS has poor convergence, especially in the presence of a correlated far-end signal.
[0037] 2) RLS has excellent performance, but has quadratic complexity in filter length.
[0038] 3) APA converges fast, but has relatively high residual echo after convergence.
[0039] Sub-band methods effectively divide the problem into different frequency bands. Then one of the three methods above can be applied within each sub-band. The Weighted Overlap-and-Add (WOLA) method falls into this category.
[0040] For the LMS algorithm, there is an important scalar parameter called the step size that controls the trade-off between convergence speed and steady-state residual echo. There are several schemes for adjusting the step size currently in use. For example, see NP-NLMS and JO-NLMS in IEEE Signal Processing Letters 13(10), 581-584 (2006), Benesty, J., Rey, H., Vega, L. R., and Tressens, S., "A nonparametric VSS NLMS algorithm," and An overview on optimized NLMS algorithms for acoustic echo cancellation, Paleologu, C, Ciochina, S., Benesty J., and Grant, S. L., Journal of EUROSPIT Signal Processing Advances, 2015:97 (2015). The idea is to use a large step size when the channel estimation error is high and the noise is low, and a small step size when the error is low and / or the noise is high. In speech-oriented applications, a voice activity detection (VAD) algorithm can be used to determine when there is a near-end sound signal present. The VAD can be fed into the step size control, making the step size equal to zero (or close to zero) during speech activity, and larger when there is no near-end speech. This is because high speech activity is expected to dominate other signals, so low or no adaptation is chosen at these times.
[0041] The APA algorithm has two parameters, a step size and a regularization parameter. Typically, the regularization parameter is set to a small fixed level to avoid numerical ill-conditioning. In principle, the step size can be controlled by methods similar to the LMS algorithm. The third parameter is the memory length, typically denoted as P. A larger value of P is beneficial for fast convergence, but a smaller P value results in lower residual echo after convergence. A method for adjusting P under different conditions is presented in Albu, F., Paleologu, C, and Benesty, J., "A variable step size evolutionary affine projection algorithm," in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 429-432, IEEE (May 2011). It appears that there is no known effective method for choosing or controlling the step size, the regularization, and P.
[0042] To address these and other issues, and as an overview, the exemplary proposals herein include the following three components, although not all components are necessary at the same time. For ease of reference, these components are labeled as C1, C2, and C3.
[0043] (C1) A new update rule, incremental maximum likelihood (IML), for adaptively learning the echo channel coefficients in a two-way audio setup. The IML update rule has two parameters: i) a fixed memory order P and ii) an adaptively set confidence parameter (CP).
[0044] (C2) A practically sensible approach to adaptively setting the CP based on the available information in an audio echo cancellation setup. This update rule can enable the IML to have fast convergence and low steady-state error, e.g., if the IML can only be run during low near-end activity (e.g., with the help of near-end speech activity detection).
[0045] (C3) An IML-based echo cancellation approach that is, e.g., robust to near-end activity. This approach involves running two IML filters in parallel. The two filters can use different assumptions to set their CPs.
[0046] Now an additional overview is given, with more detailed descriptions below.
[0047] Exemplary embodiments relate to hands-free communication in a mobile device where one (or more) loudspeaker(s) is set up to convey far-end sound to a near-end user and one (or more) microphone(s) is set up to capture near-end sound to be conveyed to the far-end user. An echo cancellation module is provided to prevent the far-end sound from propagating back to the far-end user via the chain of loudspeaker, local sound path, and microphone. An adaptive mechanism for the echo cancellation module is provided that uses a window of P past samples to form a multi-dimensional statistical model of the uncertainty in the channel estimate and updates the filter coefficients to the maximum likelihood estimate under this model. This mechanism is referred to herein as the incremental maximum likelihood (IML) algorithm. The mechanism has a control parameter, which we call the confidence parameter, that can be modified to reflect a changing balance between the level of uncertainty in the channel estimate and the power level of the near-end signal. See component (C1) above. Different embodiments differ in how the confidence parameter is modified based on available information, e.g., using components (C2, C3).
[0048] One aspect addressed by example embodiments is that the confidence parameter has a theoretically optimal value, which can be estimated by various techniques in different embodiments. In particular, analysis shows that the confidence parameter can be set, for example, equal to the ratio of the residual far-end signal power to the near-end signal power. This ratio is referred to herein as the RFNR (Residual Far-End to Near-End Ratio). Setting the confidence parameter of the IML mechanism equal to the estimated RFNR allows the adaptive mechanism to behave differently in different situations, and thus to combine some of the positive features of several other well-known adaptive mechanisms in one mechanism. For example, when the RFNR is high and the confidence parameter is set accordingly, the IML update is almost identical to the APA update, thus enabling fast reduction of high residual echo levels. When the RFNR is low and the confidence parameter is set accordingly, the IML update is almost identical to the LMS update, with a smaller step size, which is robust to high near-end noise levels and achieves low residual echo. For intermediate values of the RFNR, the IML update provides an intermediate behavior that neither the APA nor the LMS can capture alone. When the IML algorithm is operated with a confidence parameter that is approximately equal to the RFNR, example embodiments provide fast convergence and low residual echo after convergence. Like APA and LMS, the complexity of IML is only linearly related to the length of the echo cancellation filter. Note that whenever the term "equal" is used herein, this can imply substantially equal in many examples, e.g., within some (e.g., relatively small) threshold of being equal. For example, the confidence parameter can be set to be substantially equal to the RFNR within a threshold of, e.g., one percent or a few percent or less.
[0049] The RFNR is not directly observable, and various embodiments differ in the way the RFNR is estimated. By "directly observable", it is meant that the RFNR is difficult to estimate from the data, e.g., not measurable or difficult to measure. In speech-oriented applications where an accurate VAD module is available, the RFNR can be estimated based on a flat background noise model when the near-end loudspeaker is inactive, and it can simply be assumed that the RFNR is very low when near-end speech activity is detected (see component C2). In applications where an accurate VAD is not available, e.g., when the near-end signal is not just a speech signal, the confidence parameter can be effectively controlled using a pair of parallel echo cancellation filters. One filter is controlled with an aggressive estimate of the RFNR, and the other is controlled with a conservative estimate of the RFNR, both being frequently synchronized (see point C3). In example embodiments, the aggressive estimate uses a higher confidence parameter, while the conservative estimate uses a lower confidence parameter.
[0050] The technical effects of the techniques presented herein include the following.
[0051] Possible impacts of component C1 include the following. When the confidence parameter is set to be approximately equal to the ratio of the residual far-end to near-end, the IML update rule achieves, on average, a lower residual far-end signal than can be achieved with APA or NLMS, and with much lower complexity than the optimal method (e.g., RLS).
[0052] Possible impacts of components C1 and C2 together include the following. In applications where a strong near-end activity period is known or can be effectively estimated, an echo canceller working with C1 and C2 together achieves fast reduction of residual echo after a change in channel conditions (or at initialization), while also achieving very low residual echo when the channel is stable. The fast reduction is based on similarity to the reduction for APA and with faster reduction than LMS. The low residual echo is based on similar residual echo to that achieved with LMS and lower than that achieved with APA.
[0053] Possible impacts of components C1 and C3 together include the following. In general applications, an echo canceller operating with C1 and C3 achieves fast reduction of residual echo after a change in channel conditions (or at initialization), while also achieving very low residual echo when the channel is stable, while also preserving low residual echo during high near-end activity.
[0054] Having provided an overview, additional details are now provided.
[0055] Before proceeding to other details, certain concepts introduced below are characterized in mathematical form. The following table is a reference guide to parameters and their corresponding exemplary meanings:
[0056]
[0057]
[0058] This table is provided for ease of reference and is not meant to be exhaustive or limiting. Also, other names can sometimes be used to refer to these parameters.
[0059] Consider Figure 1 the settings shown. Figure 1 is a block diagram of an exemplary typical audio system with echo cancellation. There is a signal 15 from the far-end and a signal 65 to the far-end. The audio system 10 includes a loudspeaker array 12 (with three loudspeakers in this example), a microphone array 30 (with three microphones in this example), an acoustic echo canceller (AEC) 90, and a summer 76. The AEC 90 in this example includes a filter 92 with coefficients w techo canceller module 50, an adaptive weight update function 70, and a near-end activity detection module 80. Arrays 12 and 30 can have from one element to many elements, and the number of elements of each array 12, 30 need not be the same.
[0060] Signal 15 from the far-end includes a loudspeaker signal x t 11, while microphone (mic) signal y t 35 includes a noise signal z t 40, near-end signal u t 45, and far-end signal with echo 60. w * represents a channel between loudspeaker(s) 12 and microphone(s) 30. In this example, the environment of system 10 is within a room 20, and near-end signal 45 is created at least by a near-end audio source 22 such as a user (not shown).
[0061] Echo canceller 50 uses and applies coefficients w t to produce an echo estimate 75, adder 76 subtracts this echo estimate from microphone signal 35 to create an echo-canceled output e t 65. Adaptive weight update function 70 updates coefficients w t , which can also be considered weights. Near-end activity detection module 80 performs VAD and outputs a hard output (e.g., 0 for no speech detected, 1 for speech detected) or a number between (and possibly including) 0 and 1 to adaptive weight update function 70. In response thereto, adaptive weight update function 70 will or will not perform an update, e.g., using a different step size (if used).
[0062] Figure 1A is a block diagram of a communication device 110 suitable for implementing echo cancellation according to exemplary embodiments. One example of communication device 110 is a wireless device, typically a mobile device, that can access a wireless network. Communication device 110 includes one or more processors 120, one or more memories 125, one or more transceivers 130, and one or more network (N / W) interfaces (IF) 161, which are interconnected through one or more buses 127. Each of the one or more transceivers 130 includes a receiver Rx 132 and a transmitter Tx 133. The one or more buses 127 can be address, data, or control buses, and can include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, optical fiber or other optical
[0063] The communication device 110 can be wired, wireless, or both. For wireless communication, one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123. The one or more N / W I / F communicate via one or more wired links 162.
[0064] The communication device 110 includes a control module 140, which includes one or both of portions 140-1 and / or 140-2, which can be implemented in a variety of ways. The control module 140 can be implemented in hardware as control module 140-1, for example as part of the one or more processors 120. The control module 140-1 can also be implemented as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 140 can be implemented as control module 140-2, which is implemented as computer program code 123 and executed by the one or more processors 120. For example, the one or more memories 125 and computer program code 123 can be configured to, with the one or more processors 120, cause the user device 110 to perform one or more of the operations as described herein. The AEC 90 can similarly be implemented as an echo canceller module 90-1 as part of the control module 140-1, or as an echo canceller module 59-2 as part of the control module 140-2. The AEC 90 generally includes an echo canceller module 50 and an adaptive weight update function 70, and can or can not include a near-end activity detection module 80.
[0065] The computer-readable memory 125 can be of any type suitable to the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, firmware, magnetic storage devices and systems, optical storage devices and systems, fixed memory and removable memory, as non-limiting examples. The computer-readable memory 125 can be a means for storing information and / or instructions. The processor 120 can be of any type suitable to the local technical environment, and can include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multi-core processor architectures, as non-limiting examples. The processor 120 can be a means for performing functions including controlling the communication device 110 and other functions as described herein.
[0066] In general, the various embodiments of the communication device 110 can include, but are not limited to, cellular telephones such as smart phones, mobile phones, cellular phones, Voice-over-Internet Protocol (VoIP) phones, and / or wireless local loop phones, tablet computers, portable computers, indoor audio devices, immersive audio devices, vehicles or vehicle-mounted devices for, e.g., wireless V2X (vehicle-to-everything) communication, image capture devices (e.g., digital cameras), gaming devices, music storage and playback appliances, Internet appliances (including Internet- of-Things, IoT, devices), IoT devices with sensors and / or actuators for, e.g., automation applications, and portable units or terminals that incorporate combinations of such functions, laptop computers, laptop computer embedded equipment (LEEs), laptop mounted equipment (LMEs), Universal Serial Bus (USB) dongles, smart devices, wireless customer-premises equipment (CPEs), Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or wireless devices that operate in industrial and / or automation contexts), consumer electronics, devices that operate on commercial and / or industrial wireless networks, etc. That is, the communication device 110 can be any device that is capable of wireless or wireline communication.
[0067] Suppose there is a single loudspeaker (in array 12) and a single microphone (in array 30). The algorithm generalizes in a simple way to the case of multiple loudspeakers and / or multiple microphones. At time t, the AEC 90 receives as input:
[0068] 1) The latest loudspeaker signal vector: where n w is the number of coefficients in w t , and R is a set of real vectors of length n w .
[0069] 2) The latest microphone measurement: y t ∈R.
[0070] 3) The P-1 previous loudspeaker and microphone measurements: for t' = t-1,..., t-P+1, (x t′ , y t′ ).
[0071] 4) The current estimate of the echo channel coefficients w t .
[0072] Now describe how the proposed echo cancellation method, IML, updates w t as a function of its inputs.
[0073] Define n wX matrix X t = [x t ,..., x t-P+1 ], and Pxl vector Additionally, define an n w x (P-1) matrix U t-1 = [x t-1 ,..., x t-P+1 ]. (Note that U t-1 is X t without the first column.)
[0074] Given a confidence parameter c t , define the normalization factor parameter as follows:
[0075]
[0076] Then, the IML updates the coefficients w t as follows:
[0077]
[0078] Regarding the confidence parameter (CP), the only parameter in the IML description is the confidence parameter c t . First, describe how this parameter can be ideally set. Assume that the channel between the loudspeaker and the microphone can be described by , so that
[0079]
[0080] where z t represents additive Gaussian noise with variance , and u t represents the signal of the near-end user. First, consider the case that the near-end user is silent and u t = 0. In this case, the ideal situation is to set the parameter c t to Setting the parameter this way requires access to w * , which is not available. Before explaining a practical method to set the parameter c t , review two extreme cases to clarify the IML.
[0081] The first case involves the following: In this case, the misalignment error ||w t - w * || 2 dominates the additive noise, and c t should be made very large, The parameter c is set to zero. This happens at the beginning of a communication session, when the echo canceller has no reliable estimate of the coefficients. In the extreme case of c→∞, it can be shown that IML reduces to the standard APA without regularization. More generally, in this regime, c -1 acts like the regularization parameter in regularized APA.
[0082] The second case concerns the situation where the system can estimate the echo channel coefficients w * well, and additive noise dominates. In this case, c t is set to a small value, and IML reduces to LMS with small step size c.
[0083] These two extreme cases show that IML can be interpreted as an intelligent adaptive interpolation between APA and LMS, depending on the accuracy of the channel estimate and the noise level in the measurements.
[0084] Regarding the setting of the confidence parameter c, as mentioned before, ideally, IML sets the parameter c t based on the power of the misalignment at time t. But this is not available in practice. A practical alternative, inspired by the ideal choice, is to set c t to
[0085]
[0086] This practical choice works well in experiments.
[0087] Regarding the connection of IML to MLE, to understand IML and its derivatives, regularized APA (R-APA) is reviewed. R-APA updates the coefficients w t as follows:
[0088]
[0089] Here, δ denotes the regularization parameter. The n w x n w matrix P t is defined as
[0090]
[0091] Starting from w0= 0, in iteration t the following happens
[0092]
[0093] where Z t = [z t ,..., z t-P+1 ] T . Pt It is an n w - A matrix whose eigenvalues are equal to 1 and strictly less than 1. This characteristic indicates that w t The deviation in, that is, It converges to zero at an exponential rate. Therefore, approximately, w t It can be modeled as Where n w ×n w Matrix M t Indicates w t The covariance matrix. Now assume w t ~N(w o M t )and w o The maximum likelihood estimate (MLE) is as follows:
[0094]
[0095] To simplify this ML-based update rule, assume Then, the following occurs:
[0096]
[0097] in In practice, It can be approximated as To further improve this ML-based method, it's important to note that in this derivation, past P observations have been treated equally. However, the latest observation (x...) t y t ) is not used for w t New observations in the estimation. Furthermore, here, we simply assume M... t It is a diagonal matrix. To derive a better update rule, it is recommended to use IML, which uses MLE in two steps: first, estimate M based on P-1 past observations. t Secondly, the latest observations and derived estimates of M are used. t Update together w t .
[0098] Several embodiments will now be introduced and examined. In particular, the first (Embodiment 1) and the second (Embodiment 2) embodiments are described.
[0099] Example 1 addresses robustness to near-end signals via Voice Activity Detection (VAD). In some speech-oriented applications, the near-end signal is the speech signal u. tThe random process can be modeled as being either on or off. Various voice activity detection methods can be used to detect whether the near-end speech is on or off. In the case of hard VAD, the output of the voice activity detection module can be denoted as a t = 1 when speech activity is detected, and a t = 0 when no speech activity is detected, where "a" is used to indicate activity. With soft VAD, a t can take on any value between 0 (zero) and 1 (one) to reflect the estimated probability that the speech signal is in an active state.
[0100] Figure 2 is a logic flow diagram of a first embodiment, referred to as Embodiment 1. Figure 2 illustrates operations of one or more example methods, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnect devices used to perform the functions, in accordance with example embodiments. It is assumed Figure 2 The blocks in
[0101] In block 210, the communication device 110 receives input signals for one or more loudspeakers (one or more loudspeaker signals 11), one or more microphones (one or more microphone signals 35), and a VAD (via the near-end activity detection module 80). Figure 1 Then, one example embodiment works at each time step t as follows (see also
[0102] ) : Figure 2
[0103] The original confidence parameter is computed (block 220) as The voice activity detection module provides a value a t . The combined confidence parameter is computed as The updated weight vector w t+1 is computed (block 230) via an IML update step with confidence c t One example uses the convention that an update taking c t = 0 is interpreted as restricting the IML update equation to c t → 0, i.e., an LMS update with a step size of zero: w t+1 = w t This can be performed by the adaptive weight update function 70 of Figure 1 , which can also perform the weight update of the echo filter in block 240. This update modifies the response of the echo canceller module 50 accordingly.
[0104] The second embodiment, referred to as Embodiment 2, is now described. This embodiment provides robustness to near-end signals without requiring voice activity detection (VAD). That is, in some applications, a voice activity detection (VAD) unit may be unavailable or insufficient. For example, in some applications, the near-end signal 45 may be a continuous signal with variable intensity—for example, in the case of music or other ambient noise. In this case, it is advantageous to be able to track the changing channel even in the presence of a near-end signal, but an appropriate trade-off can be made between tracking speed and accuracy.
[0105] In this scenario, when the power level from the echo canceller increases, it is difficult to determine whether the increased power is due to the near-end signal u. t The increase in intensity is still due to residual echo (w) t -w * ) T x t The increase, for example due to the channel response w * The change. In the previous case, IML updates were performed with low confidence to prevent strong near-end signals from corrupting the (already accurate) weight vector w. t Ideally, in the latter case, an IML update would be performed with high confidence to correct the (inaccurate) weight vector w as quickly as possible. t That would be ideal.
[0106] Since it is difficult to distinguish between these two scenarios beforehand, this paper proposes a method in which two behavioral schemes are tried in parallel. The results of the two methods are frequently compared, at which point the appropriate action becomes clear through hindsight. Echo cancellation typically outputs the results of the low-confidence branch, but switches to the high-confidence branch when that branch demonstrates superior performance. In this way, robustness to strong near-end signals and rapid response to changes in channel response are achieved.
[0107] This method will now be described in more detail. (Turn to...) Figure 3 This figure is a block diagram of an AEC 300, referred to as Embodiment 2, according to an exemplary embodiment. In this example, the AEC 300 does not include a near-end activity detection module 80. Figure 3 The filter has two filters: a conservative filter 310 and an aggressive filter 360. The conservative filter 310 has adjustable coefficients. As indicated by reference numeral 320, and after being added to the microphone signal 35 via adder 315, an output is generated. The radical filter 360 has adjustable coefficients. As indicated by reference numeral 370, and after being added to the microphone signal 35 via adder 380, an output is generated. The conservative filter 310 has an output of adder 315, IML module 325, while aggressive filter 360 has an output of adder 380 and IML module 365. There is a controller 345, several power estimators (est) 330, 340 and 350, and a periodic synchronization block 335. Based on at least the power estimate from power estimator 350, controller 345 generates a confidence parameter for conservative filter 310 and its IML module 325 and a confidence parameter for aggressive filter 360 and its IML module 365. IML modules 325, 365 are examples of adaptive weight update function 70. Reference numerals 320 and 370 are examples of echo canceller modules 50, each producing its corresponding error output or
[0108] In short, conservative filter 310 applies filter weights (see reference numeral 320) to loudspeaker signal x t and subtracts the result from microphone signal y t IML module 325 adjusts the filter weights based on the confidence parameter provided by controller 345. Similar structure at the bottom implements aggressive filter 360, where the adjustment is based on the confidence parameter That is, aggressive filter 360 applies filter weights (see reference numeral 370) to loudspeaker signal x t and subtracts the result from microphone signal y t IML module 365 adjusts the filter weights based on the confidence parameter provided by controller 345. Periodic synchronization module 335 periodically compares the performance of the two filters, replacing the parameters of the poorer performing filter with the parameters of the better performing filter. Periodic synchronization module 335 executes the logic flow in Figure 5
[0109] In this example, two parallel echo cancellation filters 310, 360 are maintained, with the conservative filter and with the aggressive filter. The corresponding echo canceller outputs for j = 1, 2 are The output power of each filter can be computed via exponential averaging
[0110] Assuming the far-end and near-end signals are statistically independent, for fixed filter coefficients, the output power is the sum of the near-end signal power and the residual far-end echo power. Therefore, if one filter has a lower output power than the other (e.g., determined by power estimators 330, 340), then that filter must have a lower residual far-end echo, and thus the filter with the lower output power is preferred. This is an example of how to determine which of the two branches is preferred at any given time.
[0111] However, this method becomes biased and requires correction as the filter is continuously adjusted. This is because the current filter... Dependent on past observations (y τ x τ )τ<t, these observations are then usually related to the current observation (y t x t Closely related. To correct for this bias, an innovative observation is calculated, which is a transformed observation. Among them, the remote signal (Almost) the same as previous P-1 observations U t-1 =[x t-1 , ..., x t-P+1 Orthogonal.
[0112] Given the current remote signal x t The past remote signal U t-1 And (for example, a high-level) confidence parameter c, the transformed loudspeaker signal can be calculated as follows:
[0113]
[0114] The following definitions are used here:
[0115] If c -1 =0, we have This means that the transformed far-end signal is orthogonal to the most recent past far-end signal. More generally, when c -1 This is almost true when the value is very small. To obtain a hypothetical measurement... if It has already been sent, so the coefficient b t Used to generate microphone signals
[0116] The conversion process reduces the amount of time spent in the process. The statistical dependence on recent past measurements. Therefore, if the transformation error is calculated... and the corresponding transformed output power Then a more reliable measurement can be made to compare the quality of the two echo filters and .
[0117] Some notation and analysis is also needed to explain how to estimate the confidence parameter of the aggressive filter. The echo canceller output signal is as follows:
[0118]
[0119] For simplicity, assume that for some time-varying misalignment parameter m t , E[(w * -w t )(w * -w t ) T ] = m t I. If it is also assumed that w t , x t , and z t are statistically independent, then we will have the following:
[0120]
[0121] where captures the far-end signal strength, is the near-end signal strength.
[0122] Note that the output power p t and the far-end signal strength s t are empirically observable, and one wants to know m t and v t to form the ratio c t = m t / v t . Given v t , the misalignment can be computed as follows:
[0123]
[0124] and given the misalignment, the near-end signal power can be computed as follows:
[0125] v t = p t - m t s t .
[0126] For the aggressive filter 360, assume that the output power is dominated by the misalignment term, and that the near-end signal is low. To prevent infinite confidence estimates, in one exemplary embodiment the noise estimate is not allowed to go below a given fraction ε of the output power. The aggressive estimate of the misalignment is as follows:
[0127]
[0128] In this context, an example operation of the two parallel filters is described.
[0129] Example parameters used include the following:
[0130] 1) Storage length P > 2.
[0131] 2) Power averaging step size 0 < μ < 1.
[0132] 3) Minimum misalignment rate ∈ > 0.
[0133] 4) Test threshold 0 < β < 1.
[0134] 5) Multiplication factor γ > 1.
[0135] 6) Update period T.
[0136] One possible initialization procedure is as follows. At time t = 1, the two echo filters 310, 360 have the same value: Set the initial misalignment estimate to Set the initial variable s0 all to zero.
[0137] An example general update procedure is now described. The distribution is over Figure 4A and 4B Figure 4 is a logic flow diagram of the general update procedure of embodiment 2. Figure 4 illustrates the operation of one or more example methods, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnecting components used in the execution thereof, according to an example embodiment. It is assumed that the blocks in Figure 4 are executed by the communication device 110 under the control of the AEC 90 and the control module 140.
[0138] In block 405, the communication device 110 receives input signals from the loudspeaker(s) and microphone(s). The communication device 110 computes the filter outputs, for j e {1, 2}, as See blocks 410 and 425. Note that for the conservative filter 310, j = 1, and for the aggressive filter 360, j = 2. Set the echo canceller output to This sets the output to the output of the conservative filter 310. See block 415.
[0139] Update the remote signal strength See block 430.
[0140] In blocks 455 and 435, the power levels are updated for the conservative filter and aggressive filter, respectively, as follows: for j e {1, 2}, are These blocks use the power estimators 330 and 340, respectively.
[0141] In block 440, the aggressive misalignment estimate is computed as This formula is an upper bound on the misalignment given the observations; together with the following formula for c2, this yields an upper bound on the RFNR. In block 460, the conservative misalignment estimate is computed as This formula provides a low estimate of the misalignment based on the history of the low and aggressive estimates. Together with the following formula for c1, this yields a first estimate for the RFNR.
[0142] Intuitively, the reasonable low value of this ratio estimate can be based on the assumption that the misalignment (i.e., the error in estimating the coefficients) is the same as in the past. Another way to define the reasonable low value of the estimate is to be significantly lower than the aggressive estimate, e.g., 10 times lower. However, it is important to note that the confidence parameters estimates are not always different. However, they are sometimes very different, which is important for performance.
[0143] In blocks 465 and 445, the confidence parameter estimates are computed for the conservative filter and aggressive filter, respectively, for j e {1, 2}, are The term "confidence parameter" is used because this indicates the system's confidence about the measurements y t carrying useful information about the echo channel. When there is more confidence in the measurements, a larger step size can be employed. Similarly, when the confidence in the measurements is lower, a smaller step size can be employed.
[0144] The filters are updated using the IML formula with the confidence parameters c j for j e {1, 2}. This is illustrated by blocks 470 and 450, in which the weights of the conservative echo filter are updated in block 470 and the weights of the aggressive echo filter are updated in block 450.
[0145] Other possible actions include the following.
[0146] The coefficients used to form the microphone signals are computed as follows: for j e {1, 2}, are
[0147] The near-end signal strength is computed as follows: for j e {1, 2}, are
[0148] Other collateral tasks include updating the converted estimates.
[0149] 1)
[0150] 2) For j e {1,2}, compute
[0151] 3)
[0152] 4) Compute
[0153] 5) Compute
[0154] For the power level in (3), this is a running estimate of the average power level of the filtered output signal obtained by exponential averaging.
[0155] There is also a periodic update step, see Figure 5 described. Figure 5 is a logical flowchart of the periodic update rule of Example 2. Figure 5 illustrates the operation of one or more example methods, results of execution of computer program instructions embodied on a computer readable memory, functions executed by logic implemented in hardware, and / or interconnecting components used in the execution thereof, according to an example embodiment. It is assumed Figure 5 the blocks in are performed by the communication device 110 under control of the AEC 90 and the control module 140.
[0156] Block 550 indicates that after the procedure of Fig. 4 has been repeated a number of times, Figure 5 so that the error cancellation performance of the filter is repeatedly estimated. The number determines the periodicity, and examples of such are now described.
[0157] The two filters are compared and synchronized periodically, e.g. when t = kT for some update period T and any integer k. See also Figure 3 the periodic synchronization block 335 in. This procedure can have a constant factor β < 1 and a multiplication factor γ > 1.
[0158] In block 505, the conservative and aggressive echo filter coefficients, the output power and the misalignment are received in block 505. In block 510, the communication device 110 determines whether the output power of the aggressive filter is less than the output power of the conservative filter multiplied by a constant factor. If (block 510 = yes), it is considered that the aggressive filter performs better than the conservative filter.
[0159] In response, the following is performed.
[0160] The conservative filter is set This sets the coefficients (coeffs) of the conservative filter equal to the coefficients of the aggressive filter. See block 520.
[0161] Setting and setting That is, the power level of the conservative filter is set equal to the power level of the aggressive filter. See block 525.
[0162] Setting (an increased conservative estimate of misalignment). This is indicated by block 530, in which the misalignment of the conservative filter is increased by a constant multiplication factor.
[0163] Otherwise (block 510 = no), the conservative filter is deemed the best. The following is performed in response.
[0164] Setting This is illustrated by block 535, in which the coefficients of the aggressive filter are set equal to the coefficients of the conservative filter.
[0165] Setting and setting This occurs in block 540, in which the power level of the aggressive filter is set equal to the power level of the conservative filter.
[0166] In blocks 520 and 535 above, the coefficients of the poorer filter are set equal to the coefficients of the better filter. However, this is just one option. As shown in blocks 521, 536, the coefficients can instead be set to be "closer" to the other coefficients. For example, the coefficients of the poorer filter can be set to the average of the coefficients of the two filters. This will make the coefficients closer to the coefficients of the better filter, but in a more gradual manner. That is, the term "closer" can be defined as a vector norm that reduces the difference between the coefficients of the two filters (considering that the coefficients of each filter are described by a vector w).
[0167] Furthermore, while the output power is used in block 510, performance can also be used instead. See block 511. Performance can be determined as the lower output power of the echo canceller meaning better performance. There can be other performance metrics, with power output being one exemplary performance metric.
[0168] To illustrate the technical effect of the embodiment without voice activity detection, in Figure 6 and 7Results from an echo cancellation simulation are depicted. In this simulation, the far-end signal is a continuous speech signal, and the near-end signal is an intermittent speech signal plus a small amount of background noise. The echo channel from a single loudspeaker to a single microphone is a typical room acoustic echo response. The channel is generally constant, but at three different time intervals (7-9 seconds (s), 14-16 s, and 26-28 s), the channel changes significantly.
[0169] Figure 6 The intensity of the far-end signal, the near-end signal, and the residual echo as a function of time is shown. This figure illustrates the evolution of the signal power in a simulated SISO echo cancellation scenario using echo cancellation with two parallel filters running the IML algorithm. In reference number 610, the near-end signal intensity (intermittent speech plus background noise) is presented. In reference number 620, the echo signal intensity (speech far-end signal received at the near-end microphone) is presented. In reference number 630, the residual echo signal at the cancellation output is presented. Reference number 640 highlights the time intervals where the echo channel changes. Reference number 1 indicates that if the channel changes when the near-end signal is high, the tracking recovers as soon as the near-end signal decreases. Reference number 2 indicates that if the channel changes when the near-end signal is low, the tracking is fast and effective. Reference number 3 indicates that when the channel is fixed, the residual echo is not affected by the near-end signal.
[0170] Figure 7 The evolution of the normalized misalignment 20 log 10 ||w t -w * ||-20 log 10 ||w * || over time is shown. Reference numbers 1, 2, and 3 indicate the same as they do in Figure 6 In this figure, the moment when the aggressive filter is selected instead of the conservative filter is highlighted by reference number 650 - in all other cases, the conservative filter is selected. Reference number 640 highlights the time intervals where the echo channel changes. Figure 7 The evolution of the normalized misalignment 20 log Figure 6 ||w 10 ||w t -w * ||-20 log 10 ||w * || is illustrated.
[0171] During periods of low near-end signal, the misalignment and residual echo decrease rapidly - e.g. close to time 0 seconds (zero seconds) and time 17 seconds. During these periods, the algorithm correctly uses the aggressive filter running IML with high confidence parameter. When the near-end signal is strong and the channel is silent, the accuracy of the echo channel is preserved in intervals such as 1-5 seconds and 20-25 seconds. During these periods, the algorithm correctly uses the conservative filter running IML with low confidence parameter. When the channel changes, the misalignment increases temporarily. When the channel change occurs during a period of low near-end activity (cf 26-28s), the filter can adjust fast enough to keep the residual echo low. When the channel change occurs during a period of high near-end activity (see 7-9 seconds), the filter has to wait for the interruption of the near-end signal in order to learn the new echo channel.
[0172] This example illustrates how an echo canceller implementing the IML update algorithm in two parallel branches can achieve fast tracking, low residual echo, robustness to near-end signal and low complexity.
[0173] In summary, certain example embodiments can have one or more of the following advantages and technical effects.
[0174] 1) Fast convergence when the residual is high (as in APA, RLS) because IML exploits the information from the most recent P measurements when the confidence of the measurement is high.
[0175] 2) Small asymptotic residual (as in small step LMS) because IML averages the fluctuations from the near-end signal when the confidence of the measurement is low.
[0176] 3) Automatic adjustment to the near-end activity using a VAD (embodiment 1) or without a VAD (embodiment 2) due to the theoretical understanding of the optimal setting of the confidence parameter.
[0177] 4) Low computational complexity (filter length is linear) because IML uses the same computational framework as APA.
[0178] In some embodiments, the echo cancellation can be performed using a filter bank, wherein the microphone signal and the loudspeaker signal are passed through two or more parallel filters having complementary passbands to generate a plurality of sequences of subbands, wherein the echo cancellation is performed independently and in parallel within each subband, and wherein the outputs of the echo cancellation in each subband are combined to generate a final sequence of outputs in the time domain. In this case, the examples described above can be directly applied on the sequences in each subband.
[0179] In some embodiments, such as when using a filterbank based on a discrete Fourier transform (DFT), or when using a baseband representation of a carrier modulated signal, the loudspeaker and microphone signals and the estimated channel coefficients can be represented as complex values, rather than real values. The formulas given earlier extend naturally to the complex case, as will be apparent to those skilled in the art. For example, the formulas for a t (normalization factor parameter) and the paragraph t+1 The formulas for w
[0180] and
[0181]
[0182] where A H denotes the Hermitian transpose of a complex matrix or vector A.
[0183] Turning to Figure 8 , this figure is a logic flow diagram of acoustic echo cancellation using control parameters. The figure also illustrates the operation of one or more exemplary methods, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnecting components used in the execution as one exemplary embodiment. Assume that this figure is performed by the communication device 110 using the AEC 90.
[0184] In block 810, operations are performed to receive, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal. One or more loudspeakers reproduce the far-end signal.
[0185] In block 820, operations are performed to operate the adaptive echo cancellation system at least in part with at least one filter to update an estimate of a coefficient of a sound channel from the one or more loudspeakers to the one or more microphones. At least one control parameter affecting operation of the adaptive echo cancellation system is determined in block 830, the control parameter being configurable and set to at least one value of a range of values. The at least one control parameter is determined based on an accuracy of the estimate of the estimate of the coefficient of the sound channel and a characteristic of the near-end signal.
[0186] In block 840, control operations are performed by the adaptive echo cancellation system to control the at least one filter with different values of the at least one control parameter at different times.
[0187] Now give further examples.
[0188] Example 2. According to Figure 8the method of any of the examples, wherein the at least one filter comprises a first filter and a second filter, and wherein one of the different values used first on the first filter is different than another of the different values used first on the second filter.
[0189] Example 3. The method of example 2, wherein the controlling the at least one filter with different values of the at least one control parameter at different times further comprises:
[0190] controlling the first and second filters with different values of respective corresponding first and second control parameters that affect a rate of change of the respective corresponding estimate of the coefficient of the acoustic channel, wherein the value of the first control parameter set for the first filter causes the channel coefficient estimate to change at a slower rate than a rate of change caused by the value of the second control parameter for the second filter;
[0191] repeating the estimating of the error cancellation performance of the first and second filters; and
[0192] after the repeated estimating, updating coefficients of the first or second filter estimated to have lower performance to be closer to coefficients of the other of the first or second filter estimated to have higher performance.
[0193] Example 4. The method of example 3, wherein the updating further comprises updating coefficients of the first or second filter estimated to have lower performance to be equal to coefficients of the other of the first or second filter estimated to have higher performance.
[0194] Example 5. The method of example 3, wherein the performance is characterized by output power.
[0195] Example 6. The method of example 3, wherein the controlling the first and second of the at least two filters with different values of respective corresponding first and second control parameters further comprises:
[0196] determining two estimates of a residual far-to-near ratio:
[0197] the first estimate of the residual far-to-near ratio based on a past history of the first and second estimates, the first estimate being selected as a reasonably low value of the ratio;
[0198] the second estimate being selected as an upper limit of the residual far-to-near ratio based on observations of signals from the one or more microphones and the far-end signal, the second estimate being selected as a maximum value that the upper limit can be;
[0199] setting the first confidence parameter of the first adaptive filter to the first estimate;
[0200] setting the second confidence parameter of the second adaptive filter to the second estimate.
[0201] Example 7. The method of example 6, wherein the first estimate is selected to be a reasonable low value of the ratio that is significantly lower than the second estimate by a factor.
[0202] Example 8. The method of example 3, further comprising setting a power level of an estimate of the first filter or second filter estimated to have lower performance to be equal to a power level of the other of the first filter or second filter estimated to have higher performance.
[0203] Example 9. The method of example 3, further comprising increasing an estimated misalignment of the first filter in response to the second filter being estimated to have lower performance than the first filter.
[0204] Example 10. The method of example 9, wherein increasing the estimated misalignment of the first filter further comprises increasing the estimated misalignment of the first filter by a constant multiplication factor.
[0205] Example 11. The method of example Figure 8 , wherein the characteristic of the near-end signal comprises a signal strength of the near-end signal.
[0206] Example 12. The method of example 11, wherein the signal strength is characterized by an average power of the near-end signal.
[0207] Example 13. The method of example Figure 8 , wherein the determining at least one control parameter is based on estimating a ratio between an error measure in the estimate of the coefficients of the acoustic channel and a strength measure of the near-end signal.
[0208] Example 14. The method of example Figure 8 , wherein a first value of the different values used at a first time is different than a second value of the different values used at a second time.
[0209] Example 15. A computer program comprising code for performing the method of any one of examples 1 to 14 when said computer program is run on a computer.
[0210] Example 16. The computer program of example 15, wherein the computer program is a computer program product comprising a computer-readable medium bearing computer program code embodied therein for use with the computer.
[0211] Example 17. The computer program of example 15, wherein the computer program is directly loadable into the internal memory of the computer.
[0212] Example 18. An apparatus for echo cancellation of bidirectional audio communication, comprising means for:
[0213] receiving, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal, wherein one or more loudspeakers reproduce the far-end signal;
[0214] operating the adaptive echo cancellation system at least in part with at least one filter to update an estimate of a coefficient of a sound channel from the one or more loudspeakers to the one or more microphones;
[0215] determining at least one control parameter affecting operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein the determining the at least one control parameter is based on an accuracy of the estimate of the estimate of the coefficient of the sound channel and a characteristic of the near-end signal; and
[0216] controlling the at least one filter with different values of the at least one control parameter at different times by the adaptive echo cancellation system.
[0217] Example 19. The apparatus of example 15, wherein the at least one filter comprises a first filter and a second filter, and wherein one of the different values used first on the first filter is different from another of the different values used first on the second filter.
[0218] Example 20. The apparatus of example 16, wherein the controlling the at least one filter with different values of the at least one control parameter at different times further comprises:
[0219] controlling the first and second filters with different values of respective corresponding first and second control parameters affecting a rate of change of corresponding estimates of the coefficient of the sound channel, wherein the value of the first control parameter set for the first filter causes the channel coefficient estimate to change at a slower rate than a rate of change caused by the value of the second control parameter for the second filter;
[0220] repeating the estimating of the error cancellation performance of the first and second filters; and
[0221] after the repeated estimating, updating coefficients of the first or second filter estimated to have lower performance to be closer to coefficients of the other of the first or second filter estimated to have higher performance.
[0222] Example 21. The apparatus of example 17, wherein the updating further comprises updating coefficients of the first or second filter estimated to have lower performance to be equal to coefficients of the other of the first or second filter estimated to have higher performance.
[0223] Example 22. The apparatus of example 17, wherein the performance is characterized by output power.
[0224] Example 23. The apparatus of example 17, wherein controlling the first and second of the at least two filters with different values of respective corresponding first and second control parameters further comprises:
[0225] determining two estimates of a residual far-to-near ratio:
[0226] the first estimate of the residual far-to-near ratio based on a past history of the first estimate and the second estimate, the first estimate being selected as a reasonably low value of the ratio;
[0227] the second estimate being selected as an upper limit of the residual far-to-near ratio based on observations of signals from the one or more microphones and the far-end signal, the second estimate being selected as a maximum value that the upper limit can be;
[0228] setting the first confidence parameter of the first adaptive filter to the first estimate;
[0229] setting the second confidence parameter of the second adaptive filter to the second estimate.
[0230] Example 24. The apparatus of example 20, wherein the first estimate is selected as a reasonably low value of the ratio that is significantly lower than the second estimate by a factor.
[0231] Example 25. The apparatus of example 17, wherein the means are further configured to perform setting an estimated power level of the first or second filter estimated to have lower performance to be equal to a power level of the other of the first or second filter estimated to have higher performance.
[0232] Example 26. The method of example 17, wherein the means are further configured to perform increasing the estimated misalignment of the first filter in response to the second filter being estimated to have lower performance than the first filter.
[0233] Example 27. The apparatus of example 23, wherein increasing the estimated misalignment of the first filter further comprises increasing the estimated misalignment of the first filter by a constant multiplication factor.
[0234] Example 28. The apparatus of example 15, wherein the characteristic of the near-end signal comprises a signal strength of the near-end signal.
[0235] Example 29. The apparatus of example 25, wherein the signal strength is characterized by an average power of the near-end signal.
[0236] Example 30. The apparatus of example 15, wherein the determining at least one control parameter is based on estimating a ratio between an error measure in the estimates of the coefficients of the acoustic channel and a strength measure of the near-end signal.
[0237] Example 31. The apparatus of example 15, wherein a first value of the different values used at a first time is different than a second value of the different values used at a second time.
[0238] Example 32. The apparatus of any preceding apparatus example, wherein the means comprise:
[0239] at least one processor; and
[0240] at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the performance of the apparatus.
[0241] Example 33. An apparatus for echo cancellation of bidirectional audio communication, comprising:
[0242] one or more processors; and
[0243] one or more memories including computer program code,
[0244] wherein the one or more memories and the computer program code are configured to, with the one or more processors, cause the apparatus to:
[0245] receive, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal, wherein one or more loudspeakers reproduce the far-end signal;
[0246] operating the adaptive echo cancellation system, at least in part, with the at least one filter to update an estimate of a coefficient of a sound channel from the one or more loudspeakers to the one or more microphones;
[0247] determining at least one control parameter that affects operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein the determining the at least one control parameter is based on an accuracy of the estimate of estimating the coefficient of the sound channel and a characteristic of the near-end signal; and
[0248] controlling the at least one filter with different values of the at least one control parameter at different times by the adaptive echo cancellation system.
[0249] Example 34. The apparatus of example 33, wherein the at least one filter comprises a first filter and a second filter, and wherein one of the different values used first on the first filter is different than another of the different values used first on the second filter.
[0250] Example 35. The apparatus of example 34, wherein the controlling the at least one filter with different values of the at least one control parameter at different times further comprises:
[0251] controlling the first filter and second filter with different values of respective corresponding first and second control parameters that affect a rate of change of a corresponding estimate of the coefficient of the sound channel, wherein a value of the first control parameter set for the first filter causes the channel coefficient estimate to change at a slower rate than a rate of change caused by a value of the second control parameter for the second filter;
[0252] repeating the estimating of error cancellation performance of the first filter and second filter; and
[0253] after the repeating of the estimating, updating the coefficient of the first filter or second filter estimated to have lower performance to be closer to a coefficient of the other of the first filter or second filter estimated to have higher performance.
[0254] Example 36. The apparatus of example 35, wherein the updating further comprises updating the coefficient of the first filter or second filter estimated to have lower performance to be equal to the coefficient of the other of the first filter or second filter estimated to have higher performance.
[0255] Example 37. The apparatus of example 35, wherein the performance is characterized by output power.
[0256] Example 38. The apparatus of example 35, wherein controlling the first and second filters of the at least two filters with different values of respective corresponding first and second control parameters further comprises:
[0257] determining two estimates of a residual far-to-near ratio:
[0258] the first estimate of the residual far-to-near ratio based on a past history of the first estimate and the second estimate, the first estimate being selected as a reasonably low value of the ratio;
[0259] the second estimate being selected as an upper limit of the residual far-to-near ratio based on observations of signals from the one or more microphones and the far signal, the second estimate being selected as a maximum value that the upper limit can be;
[0260] setting the first confidence parameter of the first adaptive filter to the first estimate;
[0261] setting the second confidence parameter of the second adaptive filter to the second estimate.
[0262] Example 39. The apparatus of example 38, wherein the first estimate is selected as a reasonably low value of the ratio that is a factor lower than the second estimate.
[0263] Example 40. The apparatus of example 35, wherein the one or more memories and the computer program code are further configured to, with the one or more processors, cause the apparatus to set an estimated power level of the first or second filter estimated to have lower performance to be equal to the power level of the other of the first or second filter estimated to have higher performance.
[0264] Example 41. The apparatus of example 35, wherein the one or more memories and the computer program code are further configured to, with the one or more processors, cause the apparatus to increase an estimated misalignment of the first filter in response to the second filter being estimated to have lower performance than the first filter.
[0265] Example 42. The apparatus of example 41, wherein increasing the estimated misalignment of the first filter further comprises increasing the estimated misalignment of the first filter by a constant multiplication factor.
[0266] Example 43. The apparatus of example 33, wherein the characteristic of the near-end signal comprises a signal strength of the near-end signal.
[0267] Example 44. The apparatus of example 43, wherein the signal strength is characterized by an average power of the near-end signal.
[0268] Example 45. The apparatus of example 33, wherein the determining the at least one control parameter is based on a ratio between a measure of error in the estimate of the coefficients of the acoustic channel and a measure of strength of the near-end signal.
[0269] Example 46. The apparatus of example 33, wherein a first value of the different values used at a first time is different from a second value of the different values used at a second time.
[0270] Example 47. A computer program product comprising a computer readable storage medium bearing computer program code embodied therein for use with a computer, the computer program code comprising:
[0271] code for receiving, at an adaptive echo cancellation system, audio signals from one or more microphones based at least in part on a near-end signal and a reproduced far-end signal, wherein one or more loudspeakers reproduce the far-end signal;
[0272] code for operating the adaptive echo cancellation system at least in part with at least one filter to update an estimate of coefficients of an acoustic channel from the one or more loudspeakers to the one or more microphones;
[0273] code for determining at least one control parameter affecting operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein the determining the at least one control parameter is based on estimating accuracy of the estimate of the coefficients of the acoustic channel and a characteristic of the near-end signal; and
[0274] code for controlling, by the adaptive echo cancellation system, the at least one filter with different values of the at least one control parameter at different times.
[0275] As used in this application, the term“circuitry” can refer to one or more or all of the following:
[0276] (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and
[0277] (b) a combination of hardware circuitry and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuitry with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processors), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and
[0278] (c) hardware circuitry and / or processor(s), such as one or more microprocessors, or a portion of one or more microprocessors, that requires software (e.g., firmware) for operation, but software that can not be present when it is not needed for operation.
[0279] The definition of circuit applies to all uses of this term in this application, including in any claims. As another example, as used in this application, the term circuit also encompasses implementations involving only hardware circuitry or a processor (or multiple processors) or hardware circuitry or a portion of one or more processors and its (or their) accompanying software and / or firmware. For example, if applicable to a particular claim element, the term circuit also encompasses a baseband integrated circuit or processor integrated circuit for a mobile device or similar integrated circuits in a server, cellular network device, or other computing or network device.
[0280] Embodiments herein can be implemented in software (executed by one or more processors), hardware (e.g., an application specific integrated circuit), or a combination of software and hardware. In one example embodiment, software (e.g., application logic, an instruction set) is maintained in any of various conventional computer-readable media. In the context of this document, a "computer-readable medium" can be any media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer, with one example of a computer described and depicted in Figure 1A An example of a computer is depicted and described in U.S. Patent Application No. 16 / 209,407, filed December 3, 2018, entitled “Systems and Methods for Providing a Virtualized Environment for a Mobile Device,” which is incorporated by reference herein in its entirety. Computer-readable media can include a computer- readable storage medium or media (e.g., memory 125 or other device) that can be any media or means that can contain, store, and / or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer. Computer-readable storage media does not include propagating signals.
[0281] If desired, the different functions discussed herein can be performed in a different order and / or concurrently with each other. Additionally, one or more of the functions described above can be optional or can be combined.
[0282] Although various aspects of the application are set forth in the independent claims, other aspects of the application include other combinations of features from the described embodiments and / or dependent claims with the features of the independent claims, and not just the combinations explicitly set forth in the claims.
[0283] It is also noted herein that while the above describes example embodiments of the application, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications which can be made without departing from the scope of the present application as defined in the appended claims.
[0284] The following abbreviations which can occur in the specification and / or drawings, are defined as follows:
[0285] 5G Fifth Generation
[0286] AEC Acoustic Echo Cancellation or Acoustic Echo Canceller
[0287] APA Affine Projection Algorithm
[0288] cf Compare
[0289] coeffs Coefficients
[0290] CP Confidence Parameter
[0291] IML Incremental Maximum Likelihood
[0292] JO-NLMS Joint-Optimized Normalized Least Mean Square
[0293] LMS Least Mean Square
[0294] Mic Microphone
[0295] MIMO Multiple Input Multiple Output
[0296] MISO Multiple Input Single Output
[0297] MLE Maximum Likelihood Estimation
[0298] NLMS Normalized Least Mean Square
[0299] NP-NLMS Non-Parametric Normalized Least Mean Square
[0300] R-APA Regularized Affine Projection Algorithm
[0301] RFNR Residual Far-to-Near Ratio
[0302] RLS Recursive Least Squares
[0303] s Second
[0304] SISO Single Input Single Output
[0305] VAD voice activity detection
[0306] WOLA weight overlap add
Claims
1. A method for echo cancellation for bidirectional audio communication, comprising: receiving, at an adaptive echo cancellation system, audio signals from one or more microphones based on a near-end signal and a reproduced far-end signal, wherein one or more loudspeakers reproduce the far-end signal; operating the adaptive echo cancellation system with at least one filter to update an estimate of coefficients of a sound channel from the one or more loudspeakers to the one or more microphones; determining at least one control parameter that affects operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein the determining at least one control parameter is based on an accuracy of the estimate of estimating the coefficients of the sound channel and a characteristic of the near-end signal; and controlling, by the adaptive echo cancellation system, the at least one filter at different times with different values of the at least one control parameter, wherein the at least one filter produces an output signal and the adaptive echo cancellation system performs the echo cancellation by subtracting the output signal from a signal from the one or more microphones.
2. The method of claim 1, wherein, The at least one filter comprises a first filter and a second filter, and wherein one of the different values used for the first time on the first filter is different from another of the different values used for the first time on the second filter.
3. The method of claim 2, wherein, The controlling the at least one filter at different times with different values of the at least one control parameter further comprises: controlling the first filter and second filter with different values of respective corresponding first and second control parameters that affect a rate of change of a corresponding estimate of the coefficients of the sound channel, wherein a value of the first control parameter set for the first filter causes the estimate of the coefficients of the sound channel to change at a slower rate than a rate of change caused by a value of the second control parameter for the second filter; estimating an error cancellation performance of the first filter and second filter; and after the estimating, updating coefficients of the first filter or second filter estimated to have lower performance to be closer to coefficients of the other of the first filter or second filter estimated to have higher performance.
4. The method of claim 3, wherein, The updating further comprises updating coefficients of the first filter or second filter estimated to have lower performance to be equal to coefficients of the other of the first filter or second filter estimated to have higher performance.
5. The method of claim 3, wherein, Controlling the first filter and second filter with different values of respective corresponding first and second control parameters further comprises: determining two estimates of a residual far-end to near-end ratio: the first estimate of the residual far-end to near-end ratio based on a past history of the first estimate and the second estimate, the first estimate selected as a low value of the far-end to near-end ratio; the second estimate selected as an upper limit of the residual far-end to near-end ratio based on an observation of a signal from the one or more microphones and the far-end signal, the upper limit being a maximum value that the upper limit can be; setting a first confidence parameter of the first filter to the first estimate; and setting a second confidence parameter of the second filter to the second estimate.
6. The method of claim 3, further comprising at least one of: setting a power level of an estimate of the first filter or second filter estimated to have lower performance to be substantially equal to the power level of the other of the first filter or second filter estimated to have higher performance; or increasing an estimated misalignment of the first filter in response to the second filter being estimated to have lower performance than the first filter.
7. The method of claim 6, wherein, Increasing the estimated misalignment of the first filter further comprises increasing the estimated misalignment of the first filter by a constant multiplication factor.
8. The method of claim 1, wherein, The characteristic of the near-end signal comprises a signal strength of the near-end signal.
9. The method of claim 1, wherein, The determining the at least one control parameter is based on estimating a ratio between an error measure in the estimate of the coefficients of the acoustic channel and a strength measure of the near-end signal.
10. An apparatus for echo cancellation of bidirectional audio communication, comprising: one or more processors; and one or more memories storing instructions, wherein the instructions, when executed by the one or more processors, cause the apparatus at least to: receive, at an adaptive echo cancellation system, audio signals based on a near-end signal and a reproduced far-end signal from one or more microphones, wherein one or more loudspeakers reproduce the far-end signal; operate the adaptive echo cancellation system with at least one filter to update an estimate of coefficients of an acoustic channel from the one or more loudspeakers to the one or more microphones; determine at least one control parameter affecting operation of the adaptive echo cancellation system, the control parameter being configurable and set to at least one value of a range of values, wherein the determining the at least one control parameter is based on estimating accuracy of the estimate of the coefficients of the acoustic channel and a characteristic of the near-end signal; and control the at least one filter with different values of the at least one control parameter at different times by the adaptive echo cancellation system, wherein the at least one filter produces an output signal and the adaptive echo cancellation system performs the echo cancellation by subtracting the output signal from a signal from the one or more microphones.
11. The apparatus of claim 10, wherein, The at least one filter comprises a first filter and a second filter, and wherein one of the different values used first on the first filter is different from another of the different values used first on the second filter.
12. The apparatus of claim 11, wherein, The controlling the at least one filter with different values of the at least one control parameter at different times further comprises: controlling the first filter and the second filter with different values of respective corresponding first and second control parameters that affect a rate of change of the estimates of the coefficients of the channel, wherein the values of the first control parameter set for the first filter cause the estimates of the coefficients of the channel to change at a slower rate than a rate of change caused by the values of the second control parameter for the second filter; repeating the estimating of the error cancellation performance of the first filter and the second filter; and after the repeated estimating, updating the coefficients of the first filter or the second filter estimated to have lower performance to be closer to the coefficients of the other one of the first filter or the second filter estimated to have higher performance.
13. The apparatus of claim 12, wherein, The updating further includes updating the coefficients of the first filter or the second filter estimated to have lower performance to be equal to the coefficients of the other one of the first filter or the second filter estimated to have higher performance.
14. The apparatus of claim 12, wherein, Controlling the first filter and the second filter with different values of respective corresponding first and second control parameters that affect a rate of change of the estimates of the coefficients of the channel, wherein the values of the first control parameter set for the first filter cause the estimates of the coefficients of the channel to change at a slower rate than a rate of change caused by the values of the second control parameter for the second filter; determining two estimates of a residual far-to-near ratio: the first estimate of the residual far-to-near ratio based on a past history of the first estimate and the second estimate, the first estimate selected as a low value of the far-to-near ratio; the second estimate selected as an upper limit of the residual far-to-near ratio based on observations of signals from the one or more microphones and the far signal, the second estimate selected as a maximum value of the upper limit that is possible; setting a first confidence parameter of the first filter to the first estimate; setting a second confidence parameter of the second filter to the second estimate.
15. The apparatus of claim 14, wherein, The first estimate is selected as a low value of the far-to-near ratio that is significantly lower than the second estimate by a factor.
16. The apparatus of claim 12, wherein, The instructions, when executed by the one or more processors, further cause the apparatus to set a power level of an estimate of the first filter or the second filter estimated to have lower performance to be substantially equal to a power level of the other one of the first filter or the second filter estimated to have higher performance; or increase an estimated misalignment of the first filter in response to the second filter being estimated to have lower performance than the first filter.
17. The apparatus of claim 10, wherein, The characteristic of the near signal includes a signal strength of the near signal.
18. The apparatus of claim 10, wherein, The determining of the at least one control parameter is based on estimating a ratio between an error measure in the estimates of the coefficients of the channel and a strength measure of the near signal.
19. The apparatus of claim 10, wherein, A first value of the different values used at a first time is different than a second value of the different values used at a second time.
20. The apparatus of claim 10, wherein, Controlling the at least one filter with different values of the at least one control parameter that affect a rate of change of respective estimates of coefficients of the channel.
Citation Information
Patent Citations
Sound signal processing method and device, and storage medium
CN110021289A