Echo cancellation

The method addresses echo suppression challenges by using coherence metrics and audio infill to preserve audio quality in interactive scenarios, ensuring effective echo cancellation and immersive listening.

WO2026039307A1PCT designated stage Publication Date: 2026-02-19DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/041295
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-08
Filing Date
2025-08-08
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing echo cancellation techniques struggle to accurately suppress echo while preserving desired audio content, particularly in interactive audio scenarios, leading to over-suppression or under-suppression of background and ambient sounds, which affects the immersive experience.

Method used

A method and system for residual echo cancellation using coherence metrics to determine gains for attenuating residual echo, including misalignment, nonlinearity, and minimum gains, while preserving spatial information and performing audio infill to restore background content.

Benefits of technology

Effectively suppresses echo while maintaining both stationary and non-stationary audio content, enhancing the immersive and interactive experience for listeners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025041295_19022026_PF_FP_ABST
    Figure US2025041295_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for performing residual echo cancellation are provided. In some embodiments, the techniques may involve obtaining representations of: signals captured by microphones, a signal sent to speakers of the audio capture device, and an output of an adaptive echo cancellation (AEC) block of the audio capture device, wherein the AEC block performed initial echo cancellation on the signals captured by the microphones. The techniques may further involve determining coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signal sent to the speakers, and the output of the AEC block. The techniques may further involve performing residual echo cancellation based at least in part on the coherence metrics.
Need to check novelty before this filing date? Find Prior Art

Description

ECHO CANCELLATION CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 112072, filed on 14 August 2024, U.S. Patent Application No.63 / 702,931, filed on 3 October 2024, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0001] This disclosure pertains to systems, methods, and media for echo cancellation. BACKGROUND

[0002] In audio content, the presence of echo may be annoying and distracting to listeners. Accordingly, it is useful to suppress echo in audio signals when detected. However, it can be difficult to accurately detect echo conditions in order to accurately suppress echo and not over suppress desired audio content. NOTATION AND NOMENCLATURE

[0003] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.

[0004] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).

[0005] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such asubsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X − M inputs are received from an external source) may also be referred to as a decoder system.

[0006] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.SUMMARY

[0007] Techniques for performing echo cancellation are provided herein. Various embodiments described herein provide methods, systems and devices that apply the described techniques for performing the echo cancellation.

[0008] In some embodiments, a method for performing residual echo cancellation is provided. The method may involve obtaining representations of: signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and an output of an adaptive echo cancellation (AEC) block, wherein the AEC block performed initial echo cancellation on the signals captured by the microphones. The method may further involve determining coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signal sent to the speakers, and the output of the AEC block. The method may further involve performing residual echo cancellation on the output of the AEC block based at least in part on the coherence metrics to generate an enhanced audio output signal.

[0009] In some examples, determining the coherence metrics and performing the residual echo cancellation are performed by a residual echo suppression (RES) block of the audio capture device.

[0010] In some other examples, a method of performing residual echo cancellation includes applying a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block. In some examples, the misalignment gain is determined based at least in part on a product of a coherence between the signal sent to the speakers and the output of the AEC block and a coherence between the signal sent to the speakers and the signal captured by the microphones.

[0011] In still some other examples, a method of performing residual echo cancellation includes applying a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the audio capture device. In some examples, the nonlinearity gain is determined based on a difference in coherence metrics associated with two or more microphones of the audio capture device. In some examples, at least one of the two or more microphones is a microphone on a rear side of the audio capture device.

[0012] In some additional examples, a method of performing residual echo cancellation includes applying a minimum gain to be used to attenuate residual echo while preserving localcontent. In some examples, the minimum gain is determined based on a ratio of an estimated power of the local content in the output of the AEC block to a power of the output of the AEC block. In some examples, the method further involves estimating the power of the local content based on frequency bins for which an echo is determined to exist. In some examples, the frequency bins for which the echo is determined to exist is based on a comparison of the coherence metrics between the reference signal sent to the speakers and the echo signals captured by the microphones to a threshold.

[0013] In yet other examples, the method further involves performing audio infill after performing residual echo cancellation to restore stationary content in the output of the AEC block. In some examples, performing audio infill is based on an estimate of background power spectrum associated with the stationary content. In some examples, the background power spectrum is determined based on a smoothed power spectrum and a deviation correction factor, wherein the deviation correction factor corrects for echo leakage. In some examples, performing audio infill is based on an identification of frequency bins of the output of the AEC with spectral holes. In some examples, the frequency bins with spectral holes are identified by comparing a power of the output of the AEC block to a low power threshold. In some examples, performing the audio infill comprises extrapolating estimated spectral power for the frequency bins with spectral holes based on a smoothed power estimate of the output of the AEC. In some examples, the smoothed power estimate is determined in both time and frequency domains.

[0014] According to some embodiments, a device or apparatus configured to perform residual echo cancellation is provided. The device may comprise one or more processors, and one or more processor-readable media storing instructions. Execution of the stored instructions by the one or more processors may cause the performance of: obtaining representations of: signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and an output of an adaptive echo cancellation (AEC) block of the audio capture device, wherein the AEC block performed initial echo cancellation on the signals captured by the microphones; determining coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signal sent to the speakers, and the output of the AEC block; and performing residual echo cancellation based at least in part on the coherence metrics.

[0015] According to still other embodiments, a system to perform residual echo cancellation is provided. The device may comprise one or more processors, and one or more processor- readable media storing instructions. Execution of the stored instructions by the one or moreprocessors may cause the performance of: obtaining representations of: signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and an output of an adaptive echo cancellation (AEC) block of the audio capture device, wherein the AEC block performed initial echo cancellation on the signals captured by the microphones; determining coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signal sent to the speakers, and the output of the AEC block; and performing residual echo cancellation based at least in part on the coherence metrics.

[0016] According to still other embodiments, a system to perform residual echo cancellation is provided. The system may comprise at least one speaker configured to generate a speaker output signal in response to a reference signal. The system may further comprise at least one microphone configured to capture one or more echo signals that include the speaker output with one or more related echoes and responsively generate a microphone signal. The system may further comprise an input transform block configured to: obtain the reference signal; and obtain the one or more echo signals as a representation of the microphone output signal. The system may further comprise an adaptive echo cancellation (AEC) block configured to perform initial echo cancellation on the one or more echo signals based on the reference signal to responsively generate an initial echo cancellation signal. The system may further comprise a residual echo cancellation (REC) block configured to: determine coherence metrics associated with pairwise combinations of the echo signals captured by the at least one microphone, the reference signal sent to the at least one speaker, and the initial echo cancellation signal; and perform residual echo cancellation based at least in part on the coherence metrics to generate a residual echo cancelled signal. The system may further comprise an output transform block configured to generate an audio output signal based on the residual echo cancelled signal.

[0017] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read- only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.

[0018] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methodsdisclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.

[0019] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figures 1A, 1B, and 1C are diagrams illustrating an example audio capture and playback device arranged in accordance with some embodiments.

[0021] Figure 1D illustrates a block diagram of an example echo cancellation system arranged in accordance with some embodiments.

[0022] Figure 2 is a block diagram of an echo cancellation system arranged in accordance with some embodiments.

[0023] Figure 3 is a block diagram of an example implementation of a residual echo suppression (RES) block arranged in accordance with some embodiments.

[0024] Figure 4 is a flowchart of an example process for determining a misalignment gain in accordance with some embodiments.

[0025] Figure 5 is a flowchart of an example process for determining a nonlinearity gain in accordance with some embodiments.

[0026] Figure 6 is a flowchart of an example process for determining a minimum gain in accordance with some embodiments.

[0027] Figure 7A is a flowchart of an example process 700 for performing audio infill based on an estimate of the background power spectrum in accordance with some embodiments.

[0028] Figure 7B is a flowchart of an example process 750 for performing audio infill based on an identification of spectral holes in accordance with some embodiments.

[0029] Figure 8 is a flowchart of an example process 800 for performing residual echo cancellation in accordance with some embodiments.

[0030] Figure 9A shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0031] Figure 9B illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.

[0032] Figure 9C illustrates a schematic block diagram of an example CPU implemented in the device architecture of Figure 9B that may be used to implement various aspects of the present disclosure.

[0033] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF EMBODIMENTS

[0034] In audio content, the presence of echo may be annoying and distracting to listeners. Accordingly, it is useful to suppress echo in audio signals when detected. However, it can be difficult to accurately detect echo conditions in order to accurately suppress echo and not over suppress desired audio content. It may be particularly difficult to suppress echo content in situations in which the audio content is interactive. As used herein, “interactive” content refers to audio content that is rendered such that audio objects are perceived in particular spatial positions that may be rendered based on a head orientation of the listener. In such cases, while it may be possible to effectively suppress echo while preserving non-stationary content such as speech, background and / or ambient sound, which may be crucial to yielding an immersive and interactive experience, may be over-suppressed using conventional techniques for echo cancellation. Alternatively, conventional techniques may under-suppress echo in order to preserve stationary background audio content.

[0035] Briefly stated, techniques for performing residual echo cancellation are provided. In some embodiments, the techniques may involve obtaining representations of: signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and an output of an adaptive echo cancellation (AEC) block of the audio capture device, wherein the AEC block performed initial echo cancellation on the signals captured by the microphones. The techniques may further involve determining coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signal sent to the speakers, and the output of the AEC block. The techniques may further involve performing residual echo cancellation based at least in part on the coherence metrics.

[0036] In general, the signals captured by the microphones of the audio capture device may include an echo signal. The signals captured by the microphones of the audio capture device are generally referred to herein as d(n), or d. The signal sent to the speakers of the audio capture device, which is sent to the speakers by one or more processors of the audio capture device, and is accordingly known by the audio capture device, is generally referred to herein as the reference signal, and is represented herein as x(n), or x. The output of the AEC block represents the signals captured by the microphones after initial echo cancellation has been performed by the AEC block, and is represented herein as e(n), or e.

[0037] AEC works to remove main echo content. But residual echo content may remain in the output signal due to filter misalignment, and non-linearity content. In addition to perceptual audio quality, the embodiments described herein aim to take immersive experience into consideration. In particular, the embodiments described herein aim to provide improved echo cancellation by reducing residual echo after application of AEC and / or preserving spatial information among channels. Based on these analyses, a residual echo suppression (RES) module was designed to reduce residual echo content and protect local content. Embodiments of the RES module may perform residual echo suppression using one or more of: a misalignment gain part, to suppress echo arising from misalignment of the filters of the AEC, a nonlinearity gain part, to suppress echo arising from nonlinearities of the audio capture device, and / or a minimum gain part that reduces echo while preserving and / or protecting non-stationary parts of the audio such as speech and sound events. Additionally or alternatively, the RES module may include a background estimation and the audio infill part and / or an echo context information-steered spectrum inpainting part, which are configured to protect the stationary part of the audio and improve the continuity of the output.

[0038] The residual echo cancellation may be performed based on coherence metrics between signals received by microphones of an audio capture device, a signal sent to speakers of the audio capture device (i.e., the reference signal), and an output signal of the AEC. Because the reference signal, corresponding to the signal sent to and played back by the speakers is known to the audio capture device, and because any residual echo in the output signal of the AEC will have characteristics of the reference signal as the echo is due to capture by the microphones of the signal played back by the speaker, determining coherence between the microphone signal, the reference signal, and / or the output signal of the AEC may allow residual echo to be identified and / or characterized (e.g., to assess a degree of residual echo in the output signal of the AEC). Based on these coherence metrics, echo cancellation gains may be determined, which may account for misalignment of the AEC filters, nonlinearities of the audio capture device, and preserve non-stationary content of the signal. These gains may then be applied by a residual echo cancellation (RES) block of the audio capture device.

[0039] In some embodiments, to further preserve background content, audio infill may be performed to preserve background stationary content. Such background stationary content may include, e.g., background noise such as traffic, background chatter, etc., which may, when rendered and played back particularly in an interactive spatial content, yield an interactive and immersive experience for the listener. In some embodiments, audio infill may be performedbased on an estimate of the background power spectrum. Additionally or alternatively, in some embodiments, audio infill may be performed by extrapolating a power spectrum for identified spectral holes.

[0040] Using the techniques described herein, echo may be effectively suppressed while preserving both stationary and non-stationary content. The techniques described herein may yield an enhanced signal that, when rendered, may yield an interactive and immersive experience for the listener.

[0041] The techniques described herein may be implemented by one or more processors of an audio capture and playback device.

[0042] Figure 1A is a diagram illustrating an example audio capture and playback device arranged in accordance with some embodiments. In the example illustrated in Figure 1A, the audio capture device is a mobile phone 100. Front side 102 of mobile phone 100 is associated with two microphones, top microphone 106 and bottom microphone 108. Back side 104 of mobile phone 100 includes a third microphone, back microphone 110. Note that each microphone may be configured to capture audio content from a particular direction associated with the location of the microphone with respect to mobile phone 100. An audio capture and playback device may additionally include one or more speakers, such as speaker 112. Each speaker may be configured to output audio content. Note that echo may originate when audio content output by the speaker is received by the microphones of the device.

[0043] Figure 1B is a diagram illustrating a side view of mobile phone 100 that includes top microphone 106 arranged in accordance with some embodiments. Figure 1B illustrates the location of top microphone 106 from a side view perspective of mobile phone 100.

[0044] Figure 1C is a diagram illustrating a side view of mobile phone 100 that includes bottom microphone 108 arranged in accordance with some embodiments. Figure 1C illustrates the location of bottom microphone 108 form a side view perspective of mobile phone 100.

[0045] Figure 1D illustrates a block diagram of a portion of an example echo cancellation system 150 arranged in accordance with some embodiments. Echo cancellation system 150 includes a speaker 152, a microphone 154, an adaptive echo cancellation block (AEC) 156, and a summer block 158. The echo cancellation system 150 is configured to receive an input signal, represented as x(n), which is a representation of the reference signal sent to the speaker 152, andresponsively generates an enhanced signal e(n). The representation of the reference signal x(n) is provided to AEC 156. An estimated impulse response h(n) represents an impulse response associated with the echo path between the input of speaker 152 and the output of microphone 154. Microphone 154 generates a microphone signal, which is represented as d(n). AEC 156generates an output signal represented as ^^^^^^^ which represents an estimated echo signal.Microphone signal d(n) and the estimated echo signal ^^^(^^) are provided to summer block 158.Summer block 158 generates the enhanced signal e(n) by subtracting the estimated echo signal^^^(^^) from microphone signal d(n). Responsive to at least the enhanced signal e(n), an RES block212 generates a residual echo cancelled signal 213, as will be described later with reference to Figures 2, 3, 4, 5, 6, 7A, 7B, and 8.

[0046] It should be noted that there may be non-linearities and / or other distortions in playback of the reference signal by the speaker. Accordingly, the reference signal (represented as x(n) in Figure 1D) is provided to speaker 152 for playback. However, there may be distortions in the signal that is played by speaker 152 due to the non-linearities of speaker 152. The distortions due to non-linearities of speaker 152 may be accounted for in the impulse response h(n). Similarly, there may be distortions due to non-linearities of microphone 154 such that the output signal of microphone 154 (represented in Figure 1D as d(n)) differs from the audio that is captured by microphone 154, due to non-linearities of microphone 154. These distortions may also be accounted for in impulse response h(n).

[0047] AEC 156 reduces the acoustic signal (generally represented herein as y(n)) by estimating the impulse response (generally represented herein as h(n)) of the loudspeaker- microphone system (e.g., the echo path). The estimate of the impulse response is represented in Figure 1D as ℎ^(^^). As illustrated in Figure 1D, the echo is the output of the near-end signal(represented as x(n)) through the room impulse response h as modeled by:^^(^^) = ^^்(^^)ℎ

[0048] In the equation given above, h may be represented as ℎ = [ℎ ்^, ℎ^, … ℎ^ି^] , and where^^ = [^^(^^)^^(^^ − 1) … ^^(^^ − ^^ + 1)]். X(n) represents the history ofthe far-endsignal x(n), and L is the length of the echo path.

[0049] The estimated echo signal, represented as ^^^(^^) may be generated by an adaptive filter. The estimated echo signal may be a linear combination of several inputs at time index n, and may be determined by:^^^(^^) = ^^்(^^)^^(^^)

[0050] In the equation given above, W(n) represents the weight vector of the adaptive filter(e.g., ^^(^^) = [^^^(^^) ^^ ்^(^^) … ^^^(^^)] ).

[0051] The initial echo cancelled signal output by the AEC block, represented in Figure 1D as e(n), is determined by subtracting the estimated echo signal ^^^(^^) from the output signals of the microphone represented by d(n).

[0052] It should be noted that Figure 1D illustrates a single microphone and a single speaker, however, this is only for simplicity. The equations and signal paths described above may be extended to a device that utilizes one or more microphones (e.g., one, two, three, etc.) and / or one or more speakers (e.g., one, two, three, etc.).

[0053] Additionally, it should be noted that Figure 1D is described above as illustrating a microphone and a speaker which receive and / or generate representations of signals. An audio capture device may generally receive and / or generate a continuous analog signal in the time domain. As used herein, a “representation” of a given audio signal may refer to the audio signal as represented in the frequency domain and / or a digitized or sampled version of the audio signal. Accordingly, while microphone 154 is described above as generating a microphone signal represented as d(n), and speaker 152 is provided a reference signal x(n), microphone 154 and speaker 152 may generate and / or receive continuous analog time signals. Transformation between continuous signals in the time domain and the representations of the signals (e.g., a digitized version of the signal and / or a frequency domain version of the signal) may be performed with various analog-to-digital converters and / or fast Fourier transformation (FFT) blocks that are now depicted in Figure 1D for simplicity.

[0054] Techniques disclosed herein involve residual echo cancellation techniques that may be performed by a residual echo cancellation (RES) block. In general, the RES block may receive a signal (represented as e(n) in Figure 1D) that is an output of an adaptive echo cancellation (AEC) block, where the AEC block performs an initial echo cancellation on an audio signal. The RES block may implement various techniques to perform residual echo cancellation. The residual echo cancellation may be performed based on coherence metrics associated with pairwise combinations of signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and the output of the AEC block. In some embodiments, the residual echo cancellation may involve determination and / or application of amisalignment gain (as described below in connection with Figure 4), a nonlinearity gain (as described below in connection with Figure 5), and / or a minimum gain (as described below in connection with Figure 6). Additionally or alternatively, in some embodiments, after application of one or more gains, the RES block may perform background content replacement, e.g., by performing audio infill based on background spectral estimation (e.g., as described below in connection with Figure 7A) and / or by performing spectral hole-pointing (e.g., as described below in connection with Figure 7B).

[0055] Figure 2 illustrates a block diagram of an echo cancellation system 200 arranged in accordance with some embodiments. As illustrated, system 200 may be configured to receive, as input, a reference signal 202 and a microphone signal 203. System 200 may include an input transform block 204, an AEC block 210, an RES block 212, and an output transform block 214. Input transform block 204 may be configured to generate a frequency domain representation of the reference signal 202 and the microphone signal203, which may be represented as data in reference bins 206 and microphone bins 208, respectively. The data in the microphone bins 206 represent the microphone signal, d(n), while data in the reference bins 208 represent the reference signal, x(n), which is the signal that was sent to the speakers. AEC block 210 may be configured to utilize the frequency domain representation of the microphone signal and the reference signal to perform initial echo cancellation as described previously above with reference to FIG. 1D. The output of AEC block 210 (corresponding to e(n) of Figure 1D) may be provided to RES block 212. RES block 212, which is configured to perform residual echo cancellation on the output of AEC block 210 to generate a residual echo cancelled signal in the frequency domain. The residual echo cancelled signal of RES block 212 may be provided to output transform block 214, which is configured to generate a time domain representation of the signal, represented in Figure 2 as audio output signal 216.

[0056] As illustrated, Microphone signal 203 may be obtained from microphones of an audio capture device. The audio capture device may be an audio capture and playback device that both plays back audio content via speakers, and captures audio content via microphones. The audio capture device may be a mobile phone, a tablet computer, a video conferencing system, a laptop computer, a desktop computer, a game console, etc. As described above, reference signal 202, which represents the signal provided and played back by speakers of the audio capture device, may be obtained or determined based on knowledge, by the audio capture device, of the signal that was provided to the speakers. Reference signal 202 and microphone signal 203 may be transformed from the time domain to the frequency domain via input transform block 204. Inputtransform block 204 may implement a short-time Fourier transform (STFT). Based on the frequency domain representation, frequency bins associated with the signal captured by the microphones may be extracted or stored in microphone bins 206. Similarly, frequency bins associated with the signal sent to the speakers of the audio capture device, generally referred to herein as the reference signal, may be extracted or stored in reference bins 208. In some examples, the bins described herein may correspond to specially configured memory.

[0057] AEC block 210 may perform an initial echo cancellation using the frequency domain representations of the microphone signal (e.g., d(n)) and the reference signal (e.g., x(n)) to generate or produce the initial echo cancelled signal 211, generally referred to herein as e(n), as shown in Figure 1D. RES block 212 may perform residual echo cancellation using an output of AEC block 210 to generate or produce the residual echo cancelled signal 213. In some embodiments, as described below in connection with Figures 4, 5, and 6, residual echo cancellation may be performed using coherence metrics associated with the microphone signal, the reference signal, and / or the output of AEC block 210. In some embodiments, RES block 212 may additionally perform background content estimation and / or replacement to counteract over- cancellation of ambient content by RES block 212.

[0058] The residual echo cancelled signal 213 generated by RES block 212 may be provided to output transform block 214. Output transform block 214 may be configured to transform residual echo cancelled signal 213 from the frequency domain to the time domain to generate audio output signal 216. Output transform block 214 may perform the frequency domain to time domain transformation using an STFT or other signal processing technique.

[0059] Figure 3 illustrates an example implementation of RES block 212 arranged in accordance with some embodiments. As illustrated, RES block 212 may include a misalignment gain block 302, a nonlinearity gain block 304, a minimum gain block 306, and a stationary content estimation block 308, each of which are described below in more detail. Note that, as described above in connection with Figure 2, RES block 212 is configured to receive an output signal of an AEC block to perform residual echo suppression. It should be understood that although RES block 212 is illustrated as including a misalignment gain block 302, a nonlinearity gain block 304, a minimum gain block 306, and a stationary content estimation block 308, in some embodiments, any of these blocks may be omitted.

[0060] As illustrated, RES block 212 may be configured to determine a misalignment gain using the misalignment gain block 302. The misalignment gain may attenuate residual echo resulting from misalignment of one or more filters of the AEC block. Example techniques for determining the misalignment gain are shown in and described below in connection with Figure 4.

[0061] In some embodiments, RES block 212 may be configured to determine a nonlinearity gain using nonlinearity gain block 304. The nonlinearity gain may be used to attenuate residual echo resulting from nonlinearities of the audio capture device. Example techniques for determining the nonlinearity gain are shown in and described below in connection with Figure 5.

[0062] In some embodiments, RES block 212 may be configured to determine a minimum gain using minimum gain block 306. The minimum gain may be used to attenuate residual echo after echo cancellation by the AEC block while preserving local content. Example techniques for determining the minimum gain are shown in and described below in connection with Figure 6.

[0063] In some embodiments, RES block 212 may be configured to estimate stationary content using stationary content estimation block 308. Stationary content may include ambient noises, e.g., background noise (e.g., birds chirping, traffic noise, background chatter, etc.) other than transient content (e.g., speech) that may be desirable to preserve. Stationary content estimation block 308 may be configured to estimate stationary content and replace portions of stationary content that were over-suppressed as part of echo cancellation. Example techniques for estimating and replacing stationary content are shown in and described below in connection with Figures 7A and 7B.

[0064] As described above, an RES block of an audio capture device may determine and / or apply a misalignment gain. The misalignment gain may be applied to remove residual echo that is due to misalignment of one or more adaptive filters of an AEC block. In particular, the adaptive filters of the AEC block may be configured to model the echo path, however, there may be inaccuracies in the adaptive filter parameters that result in residual echo in the output signal of the AEC. The misalignment gain may be determined using coherence metrics. For example, coherence between three signals may be determined between: the original microphone signal captured by the audio capture device microphone (generally represented herein as “d”), thereference signal that was sent to the speaker (generally represented here as “x”), and “e,” which represents the output signal of the AEC. In general, coherence represents the similarity between signals.

[0065] In general, coherence represents the similarity between two signals. Coherence may be determined based on the spectral power. In some embodiments, the misalignment gain may be determined as a function of the pairwise coherence metrics of the reference signal and the AEC output (generally referred to herein as Cxe), the reference signal and the microphone signal (generally referred to herein as Cxd), and / or the microphone signal and the AEC output (generally referred to herein as Cde). In some embodiments, the misalignment gain may be determined based on an additive reciprocal of the coherence between the reference signal and the AEC output (e.g., Cxe), and / or the additive reciprocal of the coherence between the reference signal and the microphone signal (e.g., Cxd). In particular, either of these coherence metrics being relatively high may indicate a high probability of either residual echo (in the case of a relatively high Cxe) or a high probability of echo in the microphone signal (in the case of a relatively high Cxd). By using the additive reciprocal, the gain may be low in cases of high coherence, or high in cases of low coherence. By way of example, a misalignment gain may bedetermined by:^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^ = (1 − ^^^^^^) ∗ (1 − ^^^^^^)

[0066] Figure 4 is a flowchart of an example process 400 for determining a misalignment gain in accordance with some embodiments. As shown in and described above in connection with Figure 3, an RES block of an audio capture device may implement techniques to determine and / or apply the misalignment gain. In some embodiments, blocks of process 400 may be implemented by one or more processors and / or control systems of an audio capture device. An example of such a control system is control system 910 of Figure 9A.

[0067] Process 400 includes various blocks or steps as illustrated by blocks 402, 404, and 406. The blocks of process 400 may be arranged differently from what is illustrated in Figure 4. For example, some blocks may be executed in a different order than what is shown, while in other examples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 400 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. An example process 400 may commence at 402.

[0068] At 402, process 400 can obtain frequency domain representations of: the signal captured by microphones of the audio capture device (generally referred to herein as d), the signal sent to the speakers of the audio capture device (generally referred to herein as the reference signal x), and the output of the AEC (generally referred to herein as e). Note that the output of the AEC represents the microphone signal after initial echo cancellation by the AEC.

[0069] At 404, process 400 can determine coherence metrics associated with pairwise combination of the microphone signal, the speaker signal, and the AEC signal. In some embodiments, the coherence metrics may represent pairwise coherence values. The coherence metrics may be represented as Cde, Cxd, and Cxe, which represent the coherence between the microphone signal and the AEC output, the coherence between the reference signal and the microphone signal, and the coherence between the reference signal and the AEC output, respectively. The coherence metrics may be determined based on cross-power spectrum metrics, which may in turn be determined based on the power spectrum for each signal. By way of example, the power spectrum of the microphone signal, the AEC output, and the reference signal,respectively, may be determined by:^^^^ = ^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)^^^^ = ^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)^^^^ = ^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)

[0070] The cross-power spectrum may be determined by:^^^^^^ = ^^^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)^^^^^^ = ^^^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)^^^^^^ = ^^^^^^ ∗ ^^ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^)

[0071] In the equations given above, α represents a smoothing factor. In some implementations, α may be, for example, 0.9, 0.95, 0.98, etc. The coherence metrics, which mayrepresent an average coherence between pairwise signal combinations, may be determined by:^^^^^^ = (^^^^^^ ∗ ^^^^^^^^(^^^^^^)) / (^^^^ ∗ ^^^^)^^^^^^ = (^^^^^^ ∗ ^^^^^^^^(^^^^^^)) / (^^^^ ∗ ^^^^)^^^^^^ = (^^^^^^ ∗ ^^^^^^^^(^^^^^^)) / (^^^^ ∗ ^^^^)

[0072] In the equations given above, conj() represents the conjugate of a given matrix.

[0073] At 406, process 400 can determine a misalignment gain representative of an echo cancellation gain to be used to attenuate residual echo resulting from misalignment of filters of the AEC. The misalignment gain may be determined based on one or more of the coherence metrics determined at 402. For example, in some embodiments, the misalignment gain may bedetermined by:^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^ = ^^(^^^^^^,^^^^^^,^^^^^^)

[0074] As described above, in some implementations, the misalignment gain may be determined based on the additive reciprocal of Cxe and / or Cxd. For example, in someembodiments, the misalignment gain may be determined by:^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^ = (1 − ^^^^^^) ∗ (1 − ^^^^^^)

[0075] As described above, because a relatively high value of Cxe and / or Cxd represent a high probability of residual echo in the AEC output and / or echo in the microphone signal, relatively high values of Cxe and / or Cxd may translate to a relatively low value of the misalignment gain, corresponding to a more aggressive echo cancellation.

[0076] As described above, in some embodiments, a nonlinearity gain may be determined by an RES block. The nonlinearity gain may account for residual echo after initial echo cancellation by the AEC that is due to nonlinearities in the audio capture device. In some embodiments, the nonlinearity gain may be determined based on coherence metrics, e.g., as described above in connection with block 404 of Figure 4. For example, in some embodiments, the nonlinearity gain may be determined based on a difference in coherence metrics between two or more microphones of the audio capture device. In particular, in some embodiments, the difference in coherence between different microphones and the AEC output may be used to determine the nonlinearity gain. For example, in some embodiments, the two microphones may be the top microphone and the rear microphone of the audio capture device. Because the top microphone is relatively close to the ear piece speaker of the audio capture device, the echo content associated with the top microphone may be higher than the echo content in the rear microphone signal, because the rear microphone is relatively far from the ear piece speaker and / or the loudspeaker(which may be at the bottom of the audio capture device). Accordingly, the difference in coherence between the rear microphone and the AEC output and the coherence between the top microphone and the AEC output may represent residual echo due to device nonlinearities, and may be used to determine the nonlinearity gain.

[0077] Process 500 includes various blocks or steps as illustrated by blocks 502, 504, 506, and 508. The blocks of process 500 may be arranged differently from what is illustrated in Figure 5. For example, some blocks may be executed in a different order than what is shown, while in other examples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 500 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. An example process 500 may commence at 502.

[0078] At 502, process 500 can obtain frequency domain representations of: the signal captured by microphones of the audio capture device (generally referred to herein as d), the signal sent to the speakers of the audio capture device (generally referred to herein as the reference signal x), and the output of the AEC (generally referred to herein as e). Note that the output of the AEC represents the microphone signal after initial echo cancellation by the AEC.

[0079] At 504, process 500 can determine a coherence metric representing a coherence between the microphone signal and the echo signal. The coherence metric is generally represented herein as Cde. Note that obtaining the frequency domain representations of the signals at block 502 and determining the coherence between the microphone signal and the echo signal are performed as part of determining a misalignment gain (e.g., as shown in and described above in connection with blocks 402 and 404 of Figure 4). Accordingly, in some embodiments, the signals and the coherence metric Cde may be obtained once and utilized for both the misalignment gain and the nonlinearity gain. The techniques described above in connection with block 404 for determining Cde may be used to determine Cde at block 504.

[0080] Note that, in an instance in which there are N microphones, the signal d may have N channels, and correspondingly, the output of the AEC, e, may also have N channels. Accordingly, Cde may have N channels, which each channel corresponding to the coherence between a given microphone and the initial echo cancellation of the signal from the microphone as output by the AEC. At 504, process 500 may identify the coherence metrics between the microphone signal and the AEC output for individual microphones. For example, the coherencemetric for the rear microphone may be determined by selecting the channel from Cde that corresponds to the rear microphone channel. The coherence metric for the rear microphone may be represented here as Cde_rear.

[0081] At 506, process 500 may determine a difference in coherence metric between two or more microphones of the audio capture device. For example, process 500 may determine a ratio of the coherence metric of the top and / or bottom microphone signal to the corresponding AEC output signal and the coherence metric of the rear microphone signal to the corresponding AECoutput signal. For example, process 500 may determine the ratio:^^^^^^ௗ = ^^^^^^ / ^^^^^^_^^^^^^^^

[0082] In the equation given above, Cdedrepresents the difference, or ratio, between Cde and Cde for just the rear microphone signal (represented as Cde_rear). Note that Cde may refer generally to Cde for the top microphone and / or the bottom microphone. In some implementations, process 500 may scale the ratio of the top / bottom coherence to the rear coherence metric such that the ratio is within the range of 0 to 1.

[0083] At 508, process 500 can determine the nonlinearity gain based at least in part on the difference in coherence metric value, where the nonlinearity gain is used for echo cancellation of residual echo resulting from nonlinearities of the audio capture device. For example, in some embodiments, as described above, the difference in coherence metric may be a ratio of the coherence for top / bottom microphones to the coherence for the rear microphones. Note that although the example above utilizes the coherence between the microphone signal and the output of the AEC signal, in some embodiments, the nonlinearity gain may additionally or alternatively utilize the coherence between the reference signal and the AEC signal and / or the coherence between the reference signal and the microphone signal. Regardless of which coherence metrics or combination of coherence metrics are used, differences in coherence across microphones of the audio capture device may be determined to characterize audio capture device nonlinearities.

[0084] As described above, in some embodiments, an RES block may determine a minimum gain that, when applied to the output of the AEC, attenuates residual echo while preserving local content. In particular, the local content that may be preserved through application of the minimum gain may correspond to non-stationary portions of the audio signal, such as speech, sound events, etc. In some embodiments, the minimum gain may be determined based on a ratio of an estimated power of the local content in an output of the AEC block to a power of the outputof the AEC block. In some embodiments, the power of the local content may be estimated by determining frequency bins in the output signal from the AEC block that contain echo and frequency bins that do not contain echo. The power of the local content may be calculated for frequency bins where echo is determined to not exist. Note that frequency bins that contain echo may be determined based on the coherence metrics, as described above.

[0085] Process 600 includes various blocks or steps as illustrated by blocks 602, 604, 606, 608, and 610. The blocks of process 600 may be arranged differently from what is illustrated in Figure 6. For example, some blocks may be executed in a different order than what is shown, while in other examples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 600 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. An example process 600 may commence at 602.

[0086] At 602, process 600 can obtain frequency domain representations of: the signal captured by microphones of the audio capture device (generally referred to herein as d), the signal sent to the speakers of the audio capture device (generally referred to herein as the reference signal x), and the output of the AEC (generally referred to herein as e). Note that the output of the AEC represents the microphone signal after initial echo cancellation by the AEC.

[0087] At 604, process 600 can determine coherence metrics associated with pairwise combination of the microphone signal, the speaker signal, and the AEC signal. In some embodiments, the coherence metrics may represent pairwise coherence values. The coherence metrics may be represented as Cde, Cxd, and Cxe, which represent the coherence between the microphone signal and the AEC output, the coherence between the reference signal and the microphone signal, and the coherence between the reference signal and the AEC output, respectively. The coherence metrics may be determined based on cross-power spectrum metrics, which may in turn be determined based on the power spectrum for each signal, as shown in and described above in connection with block 404 of Figure 4.

[0088] Note that, because determination of a misalignment gain and / or a nonlinearity gain (as described above in connection with Figures 4 and 5) involve determination of the coherence metrics, in some implementations, obtaining the frequency domain representations (at block 602) and determining the coherence metrics (at block 604) may be performed once, and the frequencydomain representations and the coherence metrics may be utilized to determine any of the misalignment gain, the nonlinearity gain, and / or the minimum gain.

[0089] At 606, process 600 can determine frequency bins for which echo exists based on the coherence metrics. In one example, frequency bins where echo exists may be determined by comparing the coherence between the reference signal and the microphone signal (Cxd) to a threshold, where echo is determined to exist for the frequency bin if the coherence exceeds the threshold. The threshold may be 0.005, 0.01, 0.05, or any other suitable number. By way ofexample, the frequency bins for which echo exists may be determined by:^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^ = ^^^^^^ > ^^ℎ^^^^^^ℎ^^^^^^

[0090] In some implementations, process 600 may determine frequency bins where non-linear echo content exists. The frequency bins where non-linear echo content exists may be determined based on differences in coherence metrics between different microphones (e.g., a comparison of a top or bottom microphone to the rear microphone). In one example, the coherence metric ratio between the top / bottom microphone to the rear microphone for the coherence between the microphone signal and the AEC output, represented as Cde_d, as described above in connection with Figure 5, may be compared to a threshold to identify frequency bins with non-linear echo present. The threshold may be 0.005, 0.01, 0.05, 0.1, etc. In one example, the frequency binswith non-linear echo present may be determined by:^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ = ^^^^^^_^^ < ^^ℎ^^^^^^ℎ^^^^^^

[0091] Based on the frequency bins for which echo exists, the frequency bins where echo does not exist may be determined. For example, the frequency bins for which echo does not exist maybe determined by:^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^^^^^^^^^^^^^ = ~^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^

[0092] At 608, process 600 can update a power associated with the local content based on the frequency bins for which an echo exists. For example, in some embodiments, process 600 may determine that the local power is to be determined responsive to determining that the number of echo bins is relatively low (e.g., below a threshold). In some embodiments, an estimate of the local content may be updated based on the number of echo bins being below a threshold. The estimate of the local content may be updated using an attack rule or a decay rule, where the attack rule or the decay rule is selected based on a power of the AEC output signal, generallyrepresented herein as MicBinsPower. Note that an attack rule may be used in instances in which the power of the current frame is larger than the smoothed power. Conversely, a decay or release rule may be used in instances in which the power of the current frame is smaller than the smoothed power. The power of the local content may be updated based on the estimate of the local content. For example, the estimate of the local content, represented herein as MicBinsPre,may be determined by first determining if ^^^^^^(^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^) < ^^ℎ^^^^^^ℎ^^^^^^ . Afterdetermining that the echo content is sufficiently low, process 600 can estimate the local content by updating frequency bins of MicBinsPre for which echo is not present. By way of example,the local content may be updated by:^^^^^^^^^^^^^^^^^^^^(^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^^^^^^^)= ^^^^^^^^^^^^^^^^^^^^^ ∗ ^^^ + ^^ ∗ (1 − ^^^), ^^^^^^^^^^^^^^^^^^^^^^^^ < ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ∗ ^^ଶ + ^^(1 − ^^ଶ), ^^^^^^^^^^^^^^^^^^^^^^^^ ≥ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

[0093] Inimplementations, α1 may generally be larger than α2. Example values of α1 include 0.85, 0.9, 0.95, 0.98, etc. Example values of α2include 0.01, 0.05, 0.1, or the like. In the equation givenabove, MicBinsPower may be determined by:^^^^^^^^^^^^^^^^^^^^^^^^ = ^^ ∗ ^^^^^^^^(^^)

[0094] Note that in the equation given above, MicBinsPowerPre represents the power of theupdated local content. The power of the updated local content may be determined by:^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ = ^^^^^^^^^^^^^^^^^^^^ ∗ ^^^^^^^^(^^^^^^^^^^^^^^^^^^^^)^^^^^^^^^^^^^^^^^^^^^^^^ = ^^ ∗ ^^^^^^^^(^^)

[0095] In the equations given above, MicBinsPre represents the power of the signal output by the AEC (where the signal is represented herein as e). In some embodiments, the power of theAEC output signal may be determined by:^^^^^^^^^^^^^^^^^^^^^^^^ = ^^ ∗ ^^^^^^^^(^^)

[0096] At 610, process 600 may determine a minimum gain to be applied, where the minimum gain is representative of an echo cancellation gain to be used to attenuate residual echo whilepreserving local content (e.g., non-stationary content). In some implementations, the minimum gain may be determined based on a ratio of the power of the local content to a power of the signal output by the AEC (e.g., based on a ratio of MicBinsPowerPre to MicBinsPower). In some embodiments, the minimum gain may represent a scaling of the ratio. In one example, the minimum gain may be determined by: ^^^^^^^^^^^^^^ = ^^^^^^^^(^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ )

[0097] Residual echo cancellation (e.g., by application of a misalignment gain, a nonlinearity gain, and / or a minimum gain) by the RES block may lead to discontinuities in the output signal. In particular, aggressive echo cancellation may effectively remove residual echo while over suppressing signals of interest. Conversely, conservative echo cancellation may preserve signals of interest while not suppressing residual echo. The minimum gain described above in connection with Figure 6 is applied to protect non-stationary content, such as speech. However, stationary content, such as background noise and / or ambient noise, which may provide context and an immersive feel to the output signal, may be over suppressed. In some implementations, after residual echo cancellation (e.g., using any of the techniques shown in and described above in connection with Figures 4 – 6), the stationary part of the signal may be filled in during post- processing, which may improve the continuity of the output signal and generate an output signal that provides a more immersive perception to the listener. The technique of filling in the output signal in post-processing is generally referred to herein as “audio infill.”

[0098] In some embodiments, audio infill may be performed by estimating a background spectral power associated with the stationary part of the signal after residual echo cancellation. Audio infill may be performed by replacing signal for frequency bins with echo based on a comparison of the power of the signal to the background spectral power. For example, in instances in which the power of the signal is less than the background spectral power, the signal may be replaced by an estimate of the background power spectrum.

[0099] Process 700 includes various blocks or steps as illustrated by blocks 702, 704, 706, and 708. The blocks of process 700 may be arranged differently from what is illustrated in Figure 7A. For example, some blocks may be executed in a different order than what is shown, while in other examples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 700 may be omitted, replaced, and / or split intoadditional blocks without departing from the present disclosure. An example process 700 may commence at 702.

[0100] At 702, process 700 can obtain frequency domain representations of: the signal captured by microphones of the audio capture device (generally referred to herein as d), the signal sent to the speakers of the audio capture device (generally referred to herein as the reference signal x), and the output of the AEC (generally referred to herein as e). Note that the output of the AEC represents the microphone signal after initial echo cancellation by the AEC.

[0101] At 704, process 700 can determine coherence metrics associated with pairwise combination of the microphone signal, the speaker signal, and the AEC signal. In some embodiments, the coherence metrics may represent pairwise coherence values. The coherence metrics may be represented as Cde, Cxd, and Cxe, which represent the coherence between the microphone signal and the AEC output, the coherence between the reference signal and the microphone signal, and the coherence between the reference signal and the AEC output, respectively. The coherence metrics may be determined based on cross-power spectrum metrics, which may in turn be determined based on the power spectrum for each signal, as shown in and described above in connection with block 404 of Figure 4.

[0102] Note that, obtaining the frequency domain representations (at block 702) and determining the coherence metrics (at block 704) may be performed once, and the frequency domain representations and the coherence metrics may be utilized to determine any of the misalignment gain, the nonlinearity gain, and / or the minimum gain (e.g., as shown in and described above in connection with Figures 4 – 6), and additionally utilized to perform audio infill in connection with process 700.

[0103] At 706, process 700 can determine a background power spectrum associated with stationary content of the audio signal after residual echo cancellation. The audio signal after residual echo cancellation may be after applying a misalignment gain, a nonlinearity gain, and / or a minimum gain. For example, in some embodiments, process 700 can calculate a smoothing parameter, represented herein as ^^^^௧. The smoothing parameter may be determined based on frequency bins for which echo exists. In one example, the smoothing parameter may bedetermined by:^^^^௧ = ^^(^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^)^^^^௧ = ^^^^^^(^^^^^^^^^^ℎ^^^^^^^^^^^^^^) / ^^^^^^^^^^^^^^^^^^^^^^^^

[0104] Note that, in the equation given above, the frequency bins for which echo exists may be determined as described above in connection with block 606 of Figure 6.

[0105] Process 700 can determine a smoothed power spectrum using the smoothing parameter and the output signal from the AEC, represented herein as e. For example, process 700 candetermine the smoothed power spectrum by determining:^^ = ^^ ∗ ^^^^௧ + ^^ ∗ ^^^^^^^^(^^) ∗ (1 − ^^^^௧)

[0106] A deviation correction factor, generally represented herein as Bminmay be determined. In general, a deviation correction factor reduces the impact of echo leakage when the background power spectrum is estimated. Because the background power is estimated by iterative smoothing, if echo leakage persists, the background power may be over-estimated. The deviation correction factor corrects for echo leakage in the estimation of the background power. The deviation correction factor may be determined based on the coherence metrics (e.g., determined at block 702). In one example, the deviation correction factor may be the misalignment gain as determined at block 406 of Figure 4. For example, the deviation correction factor may bedetermined by:^^^^^ = (1 − ^^^^^^) ∗ (1 − ^^^^^^)

[0107] The background power spectrum, Background, may be determined based on the smoothed power spectrum and the deviation correction factor. For example, in some embodiments, the background power spectrum may be the product of a minimum smoothed power spectrum value in each frequency bin, Pmin, and the deviation correction factor. Forexample, the background power spectrum may be determined by:^^^^^^^^^^^^^^^^^^^^ = ^^^^^ ∗ ^^^^^

[0108] At 708, process 700 can perform audio infill based on a probability of echo existing in each frequency bin and the background power spectrum. For example, in some embodiments, process 700 can compare a power of the output of the AEC with the background power. Responsive to determining that the power of the AEC output is less than the background power, audio infill can be performed by replacing frequency bins of the signal with the correspondingbackground power spectrum. In some implementations, only frequency bins for which echo exists (end therefore, for which the minimum gain may have been applied) may be updated based on the background estimate. Accordingly, frequency bins that may have been subject to over cancellation of stationary content or background content may be adjusted via audio infill.

[0109] In one example, audio infill may be performed by updating the audio signal by:^^^^^^^^^^^^^^(^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^)= ^^^^^^^^^^^^^^^^^^^^^, ^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^ < ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^, ^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^ ≥ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

[0110] In some implementations, audio infill may be performed by identifying frequency bins in the output of the AEC that comprise spectral holes. Performing audio infill may involve extrapolating estimated spectral power for the frequency bins with spectral holes based on a smoothed power estimate of the output of the AEC. In some embodiments, the smoothed power estimate is determined in both the time and frequency domains.

[0111] Process 750 includes various blocks or steps as illustrated by blocks 752, 754, and 756. The blocks of process 750 may be arranged differently from what is illustrated in Figure 7B. For example, some blocks may be executed in a different order than what is shown, while in other examples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 750 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. An example process 750 may commence at 752.

[0112] At 752, process 750 can identify spectral holes in the signal output of the AEC. In some embodiments, frequency bins that have spectral holes may be identified based on frequency bins for which echo exists and a comparison of the power of the frequency bin to a threshold indicative of lower power. For example, in some embodiments, frequency bins with spectralholes may be identified by:^^^^^^^^^^^^^^^^ = ^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^ & (^^^^^^^^^^ < ^^^^^^^^^^^^^^^^^^ℎ^^^^^^ℎ^^^^^^)

[0113] At 754, process 750 can smooth the power of the spectrum in both the time and frequency domain. For example, the power, represented as power(i, k), may be smoothed in the time domain by:^^^^^^^^^^௧(^^, ^^) = ^^^^^^^^^^௧(^^ − 1, ^^) ∗ ^^ + ^^^^^^^^^^(^^, ^^) ∗ (1 − ^^)

[0114] Similarly, the power may be smoothed in the frequency domain may be smoothed by:^^^^^^^^^^^(^^, ^^) = ^^^^^^^^^^^(^^, ^^ − 1) ∗ ^^ + ^^^^^^^^^^(^^, ^^) ∗ (1 − ^^)

[0115] In

[0116] At 756, process 750 can, for the identified frequency bins, extrapolate an estimated power value based on the smoothed power. For example, in some embodiments, process 750 can determine a weighting parameter based on the bins for which echo exists. In one example,the weighting parameter may be determined by:^^^^௧ = ^^(^^^^^^^^^^^^ℎ^^^^^^^^^^^^^^)

[0117] In some embodiments, the function used to determine the weighting parameter may be determined based on a ratio or percentage of the frequency bins for which an echo exists. For example, the weighting parameter may be relatively higher in instances with a higher ratio or percentage of frequency bins with an echo relative to a lower ratio or percentage. In one example, the weighting parameter may be determined by: ^^ 0.2, ^^^^^^ ^^^^ℎ^^^^^^^^^^^^^^^^^^^^ < 0.5^^௧ = ^ ^^^^^^ ^^^^ℎ^^^^^^^^^^^^^^^^^^^^ ≥ 0.5

[0118] Process 750based on the smoothed power.For example, the power estimate for channel (i, k) may be determined by:^^^^^^^^^^(^^, ^^) = ^^^^^^^^^^௧(^^, ^^) ∗ ^^^^௧ + ^^^^^^^^^^^(^^, ^^) ∗ (1 − ^^^^௧)

[0119] TheFor example, the power estimate may be extrapolated by: ^^^^^^(^^^^^^^^^^ )^^^^^^^^^^ = ^^^^^^^^^^ ∗ ^^ + ௧ ∗ ^^^^^^^^^^ᇱ ^^) ∗ − ^^)

[0120] or as and 806. The blocks of process 800 may be arranged differently from what is illustrated in Figure 8. For example, some blocks may be executed in a different order than what is shown, while in otherexamples, two or more blocks may be executed substantially in parallel. In some other implementations, one or more blocks of process 800 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. An example process 800 may commence at 802.

[0121] At 802, process 800 can obtain representations of: signals captured by microphones of an audio capture device, a signal sent to speakers of the audio capture device, and an output of an AEC block of the audio capture device, where the AEC block performed initial echo cancellation of the signals captured by the microphone. As described above, the signals captured by the microphones is generally represented as d, the signal sent to the speakers is generally represented as x, and the output of the AEC block is generally represented as e.

[0122] At 804, process 800 can determine coherence metrics associated with pairwise combinations of the signals captured by the microphones, the signals sent to the speakers, and the output of the AEC block. The coherence of the microphone signal and the speaker signal is generally referred to as Cxd, the coherence of the microphone signal and the AEC output is generally represented as Cde, and the coherence of the speaker signal and the AEC output is represented as Cxe. Techniques for determining each of Cxd, Cxe, and Cde are described above in connection with block 404 of Figure 4. In some implementations, the coherence metrics may include a ratio of coherences as determined using signals from different microphones (e.g., a ratio of coherence obtained using the top and / or bottom microphones to the coherence obtained using the rear microphone). Considering and comparing coherence metrics for different microphones may allow nonlinearities of the audio capture device in the initial echo cancellation to be considered.

[0123] At 806, process 800 can perform residual echo cancellation based at least in part on the coherence metrics. For example, in some embodiments, process 800 can perform residual echo cancellation by determining and applying a misalignment gain that attenuates residual echo resulting from misalignment of filters of the AEC, as shown in and described above in connection with Figure 4. As another example, in some embodiments, process 800 can perform residual echo cancellation by determining and applying a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the audio capture device, as shown in and described above in connection with Figure 5. As yet another example, in some embodiments, process 800 can perform residual echo cancellation by determining and applying a minimum gain that attenuates residual echo while preserving non-stationary content, as shown in and described above inconnection with Figure 6. Note that, in some embodiments, any combination of application of a misalignment gain, a nonlinearity gain, and a minimum gain may be used.

[0124] In some implementations, process 800 may perform audio infill to preserve background content and / or stationary content to compensate for aggressive echo cancellation. For example, audio infill may be performed based on an estimate of the background power, as shown in and described above in connection with Figure 7A. As another example, in some embodiments, audio infill may be performed based on power extrapolation with respect to identified spectral holes, as shown in and described above in connection with Figure 7B.

[0125] The output of process 800 may be an enhanced signal, e.g., for which residual echo cancellation and optionally audio infill has been performed. The enhanced signal may be saved by the audio capture device (e.g., in memory of the audio capture device), played back by the audio capture device, transmitted to another device (e.g., an audio playback device, a remote server or cloud service, etc.) for storage and / or playback, or the like.

[0126] Figure 9A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 9A are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 900 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 900 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0127] According to some alternative implementations the apparatus 900 may be, or may include, a server. In some such examples, the apparatus 900 may be, or may include, an encoder. Accordingly, in some instances the apparatus 900 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 900 may be a device that is configured for use in “the cloud,” e.g., a server.

[0128] In this example, the apparatus 900 includes an interface system 905 and a control system 910. The interface system 905 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environmentmay be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 905 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 900 is executing.

[0129] The interface system 905 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0130] The interface system 905 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 905 may include one or more wireless interfaces. The interface system 905 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 905 may include one or more interfaces between the control system 910 and a memory system, such as the optional memory system 915 shown in Figure 9A. However, the control system 910 may include a memory system in some instances. The interface system 905 may, in some implementations, be configured for receiving input from one or more microphones in an environment.

[0131] The control system 910 may, for example, include a general purpose single- or multi- chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0132] In some implementations, the control system 910 may reside in more than one device. For example, in some implementations a portion of the control system 910 may reside in a device within one of the environments depicted herein and another portion of the control system 910 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 910may reside in a device within one environment and another portion of the control system 910 may reside in one or more other devices of the environment. For example, a portion of the control system 910 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 910 may reside in another device that is implementing the cloud-based service, such as another server, a memory device, etc. The interface system 905 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.

[0133] In some implementations, the control system 910 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 910 may be configured for implementing methods for performing residual echo cancellation, or the like.

[0134] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 915 shown in Figure 9A and / or in the control system 910. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non- transitory media having software stored thereon. The software may, for example, filter audio signals to reduce noise, determine filter coefficients to perform such filtering, etc. The software may, for example, be executable by one or more components of a control system such as the control system 910 of Figure 9A.

[0135] In some examples, the apparatus 900 may include the optional microphone system 920 shown in Figure 9A. The optional microphone system 920 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 900 may not include a microphone system 920. However, in some such implementations the apparatus 900 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 910. In some such implementations, a cloud-based implementation of the apparatus 900 may be configured to receive microphone data, or a noise metric corresponding at least in part to themicrophone data, from one or more microphones in an audio environment via the interface system 910.

[0136] According to some implementations, the apparatus 900 may include the optional loudspeaker system 925 shown in Figure 9A. The optional loudspeaker system 925 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 900 may not include a loudspeaker system 925. In some implementations, the apparatus 900 may include headphones. Headphones may be connected or coupled to the apparatus 900 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).

[0137] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0138] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one ormore examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.

[0139] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.

[0140] Figure 9B illustrates a schematic block diagram of an example device architecture 901 (in this example, an apparatus 901) that may be used to implement various aspects of the present disclosure. The apparatus 901 of Figure 9B is an instance of the apparatus 900 of Figure 9A. Architecture 901 includes but is not limited to servers and client devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of any of Figures 4-8. As shown, the architecture 901 includes central processing unit (CPU) 941, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 942 or a program loaded from, for example, storage unit 948 to random access memory (RAM) 943. The CPU 941 may be, for example, an electronic processor 941. In these examples, the CPU 941 is an instance of the control system 910 of Figure 9A and the ROM 942 and RAM 943 are instances of the memory system 915. In RAM 943, the data required when CPU 941 performs the various processes is also stored, as required. CPU 941, ROM 942, and RAM 943 are connected to one another via bus 944. Input / output (I / O) interface 945 is also connected to bus 944. The bus 944 and the I / O) interface 945 are instances of the interface system 905 of Figure 9A.

[0141] The following components are connected to I / O interface 945: input unit 946, that may include a keyboard, a mouse, or the like; output unit 947 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 948 including a hard disk, or another suitable storage device; and communication unit 949 including a network interface card such as a network card (e.g., wired or wireless).

[0142] In some implementations, input unit 946 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0143] In some implementations, output unit 947 include systems with various number of speakers. Output unit 947 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0144] In some embodiments, communication unit 949 is configured to communicate with other devices (e.g., via a network). Drive 950 is also connected to I / O interface 945, as required. Removable medium 951, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 950, so that a computer program read therefrom is installed into storage unit 948, as required. A person skilled in the art would understand that although apparatus 901 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0145] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 949, and / or installed from the removable medium 951, as shown in Figure 9B.

[0146] Figure 9C illustrates a schematic block diagram of an example CPU 941 implemented in the device architecture 901 of Figure 9B that may be used to implement various aspects of the present disclosure. The CPU 941 includes an electronic processor 960 and a memory 961. The electronic processor 960 is electrically and / or communicatively connected to the memory 961 for bidirectional communication. The memory 961 may store echo cancellation software 962 and / or residual echo cancellation software 963. The memory 961 may be, for example, a ROM, a RAM, or another non-transitory computer readable medium. The electronic processor 960 may implement the echo cancellation software 962 stored in the memory 961 to perform, among other things, techniques for performing initial echo cancellation (e.g., as performed by an AEC block). Additionally, the electronic processor 960 may implement the residual echo cancellation software 963 stored in the memory 961 to perform, among other things, the methods that are described with reference to any or all of Figures 4-8 (e.g., as implemented by an REC block as described herein).

[0147] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic orany combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 941 in combination with other components of Figure 9B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0148] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0149] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0150] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer,or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0151] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):

[0152] EEE 1. A method, comprising: obtaining representations of: echo signals, d(n), captured by microphones of an audio capture device (100), a reference signal, x(n), sent to speakers of the audio capture device, and an output, e(n), of an adaptive echo cancellation (AEC) block (210), wherein the AEC block performed initial echo cancellation on the signals captured by the microphones (802); determining coherence metrics associated with: pairwise combinations of the echo signals captured by the microphones, the reference signal sent to the speakers, and the output of the AEC block (804); and performing residual echo cancellation on the output of the AEC block based at least in part on the coherence metrics (806) to generate an enhanced audio output signal.

[0153] EEE 2. The method of EEE 1, wherein determining the coherence metrics and performing the residual echo cancellation are performed by a residual echo suppression (RES) block.

[0154] EEE 3. The method of any one of EEEs 1 or 2, wherein performing residual echo cancellation comprises applying a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block.

[0155] EEE 4. The method of EEE 3, wherein the misalignment gain is determined based at least in part on a product of a coherence between the reference signal sent to the speakers and the output of the AEC block and a coherence between the reference signal sent to the speakers and the echo signals captured by the microphones.

[0156] EEE 5. The method of any one of EEE 1-4, wherein performing residual echo cancellation comprises applying a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the audio capture device.

[0157] EEE 6. The method of EEE 5, wherein the nonlinearity gain is determined based on a difference in coherence metrics associated with two or more microphones of the audio capture device.

[0158] EEE 7. The method of EEE 6, wherein at least one of the two or more microphones is a microphone on a rear side of the audio capture device.

[0159] EEE 8. The method of any one of EEEs 1-7, wherein performing residual echo cancellation comprises applying a minimum gain to be used to attenuate residual echo while preserving local content.

[0160] EEE 9. The method of EEE 8, wherein the minimum gain is determined based on a ratio of an estimated power of the local content in the output of the AEC block to a power of the output of the AEC block.

[0161] EEE 10. The method of EEE 9, further comprising estimating the power of the local content based on frequency bins for which an echo is determined to exist.

[0162] EEE 11. The method of EEE 10, wherein the frequency bins for which the echo is determined to exist is based on a comparison of the coherence metrics between the reference signal sent to the speakers and the echo signals captured by the microphones to a threshold.

[0163] EEE 12. The method of any one of EEEs 1-11, further comprising performing audio infill after performing residual echo cancellation to restore stationary content in the output of the AEC block.

[0164] EEE 13. The method of EEE 12, wherein performing audio infill is based on an estimate of background power spectrum associated with the stationary content.

[0165] EEE 14. The method of EEE 13, wherein the background power spectrum is determined based on a smoothed power spectrum and a deviation correction factor, wherein the deviation correction factor corrects for echo leakage.

[0166] EEE 15. The method of EEE 9, wherein performing audio infill is based on an identification of frequency bins of the output of the AEC block with spectral holes.

[0167] EEE 16. The method of EEE 15, wherein the frequency bins with spectral holes are identified by comparing a power of the output of the AEC block to a low power threshold.

[0168] EEE 17. The method of any one of EEEs 15 or 16, wherein performing the audio infill comprises extrapolating estimated spectral power for the frequency bins with spectral holes based on a smoothed power estimate of the output of the AEC block.

[0169] EEE 18. The method of EEE 17, wherein the smoothed power estimate is determined in both time and frequency domains.

[0170] EEE 19. An apparatus configured for implementing the method of any one of EEEs 1- 18.

[0171] EEE 20. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of EEEs 1-18.

[0172] EEE 21. A device, comprising: one or more processors (941); and one or more processor-readable media (961) storing instructions which, when executed by the one or more processors, cause of the device to: obtain representations of: one or more echo signals, d(n), captured by one or more microphones (154), reference signals, x(n), sent to one or more speakers (152), and an enhanced signal, e(n), output by an adaptive echo cancellation (AEC) block (210), wherein the AEC block (210) performed initial echo cancellation on the echo signals, d(n), captured by the one or more microphones (154); determine coherence metrics associated with pairwise combinations of the echo signals, d(n), captured by the one or more microphones, the reference signals, x(n), sent to the one or more speakers, and the enhanced signal, e(n), of the AEC block (804); perform residual echo cancellation based at least in part on the coherence metrics (806) for an echo cancelled signal (213); and generate an audio output signal (216) based on the echo cancelled signal (213).

[0173] EEE 22. The device of EEE 21, wherein the residual echo cancellation applies a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block.

[0174] EEE 23. The device of any one of EEEs 21 or 22, wherein the residual echo cancellation applies a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the device.

[0175] EEE 24. The device of any one of EEEs 21-23, wherein the residual echo cancellation applies a minimum gain to be used to attenuate residual echo while preserving local content.

[0176] EEE 25. A system, comprising: at least one speaker (152) configured to generate a speaker output signal in response to a reference signal x(n) (202); at least one microphone (154) configured to capture one or more echo signals that include the speaker output with one or more related echoes and responsively generate a microphone signal (203); an input transform block (204) configured to: obtain the reference signal, x(n); obtain the one or more echo signals, d(n), as a representation of the microphone output signal (203); an adaptive echo cancellation (AEC) block (210) configured to: perform initial echo cancellation on the one or more echo signals, d(n), based on the reference signal, x(n), to responsively generate an initial echo cancellation signal e(n) (211); a residual echo cancellation (REC) block (212) configured to: determine coherence metrics (804) associated with pairwise combinations of the echo signals captured by the at least one microphone (154), the reference signal sent to the at least one speaker (152), and the initial echo cancellation signal (211); perform residual echo cancellation based at least in part on the coherence metrics (806) to generate a residual echo cancelled signal (213); and an output transform block (204) configured to generate an audio output signal (216) based on the residual echo cancelled signal (213).

[0177] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.

Claims

CLAIMS 1. A method, comprising: obtaining representations of: echo signals, d(n), captured by microphones of an audio capture device (100), a reference signal, x(n), sent to speakers of the audio capture device, and an output, e(n), of an adaptive echo cancellation (AEC) block (210), wherein the AEC block performed initial echo cancellation on the signals captured by the microphones (802); determining coherence metrics associated with: pairwise combinations of the echo signals captured by the microphones, the reference signal sent to the speakers, and the output of the AEC block (804); and performing residual echo cancellation on the output of the AEC block based at least in part on the coherence metrics (806) to generate an enhanced audio output signal.

2. The method of claim 1, wherein the echo signals, ^^(^^) are captured by two or more microphones of the audio capture device.

3. The method of claim 2, wherein at least one of the two or more microphones is a microphone a rear side of the audio capture device.

4. The method of any one of claims 1 to 3, wherein the coherence metrics are determined and the residual echo cancellation is performed by a residual echo suppression (RES) block.

5. The method of any one of claims 1 to 4, wherein performing residual echo cancellation comprises applying a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block, wherein the misalignment gain is determined as a function of the coherence metrics (806).

6. The method of claim 5, wherein the misalignment gain is determined based at least in part on a product of a coherence between the reference signal sent to the speakers and the output of the AEC block and a coherence between the reference signal sent to the speakers and the echo signals captured by the microphones.

7. The method of any one of claims 2 to 6, wherein performing residual echo cancellation comprises applying a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the audio capture device, wherein the nonlinearity gain is determined based on the coherence metrics (806) associated with the two or more microphones of the audio capture device.

8. The method of claim 6, wherein the nonlinearity gain is determined based on a difference in coherence metrics associated with the two or more microphones of the audio capture device.

9. The method of any one of the preceding claims, wherein performing residual echo cancellation comprises applying a minimum gain to be used to attenuate residual echo while preserving local content, wherein the minimum gain is determined based on the ratio of an estimated power of the local content in the output of the AEC block to a power of the output of the AEC block, wherein the local content corresponds to non-stationary portions of the audio signal.

10. The method of claim 8, further comprising estimating the power of the local content based on frequency bins for which an echo is determined to exist.

11. The method of claim 10, wherein the frequency bins for which the echo is determined to exist is based on a comparison of the coherence metrics between the reference signal sent to the speakers and the echo signals captured by the microphones to a threshold.

12. The method of any one of claims 1-10, further comprising performing audio infill after performing residual echo cancellation to restore stationary content in the output of the AEC block.

13. The method of claim 12, wherein performing audio infill is based on an estimate of background power spectrum associated with the stationary content.

14. The method of claim 13, wherein the background power spectrum is determined based on a smoothed power spectrum and a deviation correction factor, wherein the deviation correction factor corrects for echo leakage.

15. The method of claim 12, wherein performing audio infill is based on an identification of frequency bins of the output of the AEC block with spectral holes.

16. The method of claim 15, wherein the frequency bins with spectral holes are identified by comparing a power of the output of the AEC block to a low power threshold.

17. The method of any one of claims 15 or 16, wherein performing the audio infill comprises extrapolating estimated spectral power for the frequency bins with spectral holes based on a smoothed power estimate of the output of the AEC block.

18. The method of claim 17, wherein the smoothed power estimate is determined in both time and frequency domains.

19. An apparatus configured for implementing the method of any one of EEEs 1-18.

20. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of EEEs 1-18.

21. A device, comprising: one or more processors (941); and one or more processor- readable media (961) storing instructions which, when executed by the one or more processors, cause the device to: obtain representation of: one or more echo signals, d(n), captured by one or more microphones (154), reference signals, x(n), sent to one or more speakers (152), and an enhanced signal, e(n), output by an adaptive echo cancellation (AEC) block (210), wherein the AEC block (210) performed initial echo cancellation on the echo signals, d(n), captured by the one or more microphones (154);determine coherence metrics associated with pairwise combinations of the echo signals, d(n), captured by the one or more microphones, the reference signals, x(n), sent to the one or more speakers, and the enhanced signal, e(n), of the AEC block (804); perform residual echo cancellation based at least in part on the coherence metrics (806) for an echo cancelled signals (213); and generate an audio output signal (216) based on the echo cancelled signal (213).

22. The device of claim 21, wherein the residual echo cancellation applies a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block, wherein the misalignment gain is determined as a function of the coherence metrics.

23. The device of any one of claims 21 or 22, wherein the residual echo cancellation applies a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the device, wherein the nonlinearity gain is determined as a function of the coherence metrics associated with the two or more microphones of the audio capture device.

24. The device of any one of claims 21 to 23, wherein the residual echo cancellation applies a minimum gain to be used to attenuate residual echo while preserving local content, wherein the local content corresponds to non-stationary portions of the audio signal.

25. A system, comprising: at least one speaker (152) configured to generate a speaker output signal in response to a reference signal ^^(^^) (202); at least one microphone (154) configured to capture one or more echo signals that include the speaker output with one or more related echoes and responsively generate a microphone signal (203); an input transform block (204) configured to: perform initial echo cancellation on the one or more echo signals, ^^(^^), based on the reference signal, ^^(^^), to responsively generate an initial echo cancellation signal ^^(^^) (211);a residual echo cancellation (REC) block (212) configured to: determine coherence metrics (804) associated with pairwise combinations of the echo signals captured by the at least one microphone (154), the reference signal sent to the at least one speaker (152), and the initial echo cancellation signal (211); perform residual echo cancellation based at least in part on the coherence metrics (806) to generate a residual echo cancelled signals (213); and an output transform block (204) configured to generate an audio output signal (216) based on the residual echo cancelled signal (213).

26. The system of claim 25, wherein the residual echo cancellation block applies a misalignment gain that attenuates residual echo resulting from misalignment of one or more filters of the AEC block, wherein the misalignment gain is determined as a function of the coherence metrics.

27. The system of any one of claims 25 or 26, wherein the residual echo cancellation block applies a nonlinearity gain that attenuates residual echo resulting from nonlinearities of the device, wherein the nonlinearity gain is determined as a function of the coherence metrics associated with the two or more microphones of the audio capture device.

28. The system of any one of claims 25 to 27, wherein the residual echo cancellation applies a minimum gain to be used to attenuate residual echo while preserving local content, wherein the local content corresponds to non-stationary portions of the audio signal.

Citation Information

Patent Citations

  • Non-linear echo cancellation

    US20130216056A1

  • Acoustic Echo Suppression

    US20150181018A1