State detection for acoustic echo cancellation

A dual-path filter model with a state detector and adaptive filters addresses suboptimal performance in acoustic echo cancellers by identifying states and adjusting filter operations, enhancing echo cancellation efficiency.

WO2025217199A1PCT designated stage Publication Date: 2025-10-16DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023714
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-30
Filing Date
2025-04-08
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing acoustic echo cancellers in subband domains face challenges in adapting to varying acoustic conditions, leading to suboptimal performance and echo removal efficiency due to issues like double talk and echo path changes.

Method used

Implementing a dual-path filter model with a main and shadow adaptive filter, utilizing a state detector to identify converging, double talk, and stable states, and adjusting the adaptive filter's step size and operation accordingly.

Benefits of technology

Improves the performance of acoustic echo cancellation by enhancing the adaptation rate and efficiency of adaptive filters, particularly during converging and double talk states, leading to better echo removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000007_0001
    Figure IMGF000007_0001
  • Figure IMGF000013_0001
    Figure IMGF000013_0001
  • Figure IMGF000013_0002
    Figure IMGF000013_0002
Patent Text Reader

Abstract

Systems, devices and methods for acoustic echo cancellation are described. One example method includes performing a first filtering operation on an audio input signal received from an audio input device, determining, based on one or more of a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state, performing, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal, and adjusting, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.
Need to check novelty before this filing date? Find Prior Art

Description

STATE DETECTION FOR ACOUSTIC ECHO CANCELLATION CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 086880, filed on 9 April 2024, U.S. Provisional Application No. 63 / 655,941, filed on 4 June 2024, and European Application No. 24203655.6 filed on 30 September 2024, each of which is incorporated by reference herein in its entirety. FIELD OF THE DISCLOSURE

[0002] Various example embodiments relate generally to acoustic echo cancellation. In particular, example embodiments are directed to a system, method, or computer program product configured for state event detection based on a dual-path filter model for acoustic echo cancellation. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS

[0003] Acoustic echo cancellers are often implemented in the subband domain for both performance and cost reasons. A subband domain acoustic echo canceller normally includes a subband acoustic echo canceller for each of a plurality of subbands. Furthermore, each subband acoustic echo canceller normally implements multiple adaptive filters, each of which is optimal in different acoustic conditions. The multiple adaptive filters are controlled by adaptive filter management modules that operate according to heuristics, so that the overall subband acoustic echo canceller may have the best characteristics of each filter.

[0004] Disclosed herein are various embodiments for adapting the performance of an acoustic echo cancellation system based on detected features associated with the performance of the adaptive filters. For example, the adaptive filters estimate an acoustic signal detected by an audio input device. The estimated acoustic signals are compared to an actual acoustic signal to determine an error of each adaptive filter. The error values, as well as the power of the audio input signal, are used to determine a state of the acoustic echo cancellation system. The state of the acoustic echo cancellation system is then used to alter operation of the adaptive filters, such as by applying the adaptive filters multiple times or by adjusting a step size associated with an adaptation rate of the adaptive filters. Such implementations have the advantages of improved performance of theadaptive filters, increases the adaptive rate of the filters, and an increased rate of echo removal.

[0005] According to an example embodiment, provided is an audio processing method comprising performing a first filtering operation on an audio input signal received from an audio input device, determining, based on a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state, performing, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal, and adjusting, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

[0006] According to another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising performing a first filtering operation on an audio input signal received from an audio input device, determining, based on a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state, performing, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal, and adjusting, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

[0007] According to yet another example embodiment, provided is an apparatus for performing acoustic echo cancellation, the apparatus comprising an input device configured to receive an audio input signal and an electronic processor connected to the input device. The electronic processor is configured to perform a first filtering operation on the audio input signal received from the audio input device, determine, based on a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state, perform, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal, and adjust, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of thefirst adaptive filter and the second adaptive filter is altered. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0009] FIG. 1 depicts a diagram of an example echo cancellation system.

[0010] FIG. 2 depicts a diagram of another example echo cancellation system.

[0011] FIG. 3 depicts a plot indicating an error difference between a first adaptive filter type and a second adaptive filter type over a plurality of frames.

[0012] FIG. 4A depicts a block diagram of an example electronic device architecture suitable for implementing example embodiments of the present disclosure.

[0013] FIG. 4B depicts a block diagram of an example controller.

[0014] FIG. 5 depicts a block diagram of an example state detection system.

[0015] FIG. 6 depicts a block diagram of an example acoustic echo cancellation algorithm.

[0016] FIG. 7 depicts a block diagram for an example method of determining a state of the acoustic echo cancellation system of FIG. 2.

[0017] FIG. 8 depicts a block diagram of an acoustic echo cancelling algorithm with echo re- filtering.

[0018] FIG. 9 depicts a graph of adaptive filter step size.

[0019] FIG. 10 depicts a block diagram of an example method of controlling the acoustic echo cancellation system of FIG. 2. DETAILED DESCRIPTION

[0020] This disclosure and aspects thereof can be embodied in various forms, including hardware, devices or circuits controlled by computer-implemented methods, computer programproducts, computer systems and networks, user interfaces, and application programming interfaces; as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and the like. The foregoing is intended solely to give a general idea of various aspects of the present disclosure and does not limit the scope of the disclosure in any way.

[0021] Acoustic echo is a major audio impairment impacting real-time audio / video communications, such as voice call and video call. Acoustic echo has significant impact on the duplex interactive audio experience. The performance of echo management is dependent on the convergence rate of adaptive filters, double talk conditions, and echo path change (for example, human behavior, occlusion, device movement, and the like). Accordingly, cancelling acoustic echo is desirable.

[0022] FIG. 1 illustrates a diagram of an example echo cancellation system 100. The example echo cancellation system 100 includes an audio input device 102 (e.g., a microphone) that provides an original microphone signal d(n), an audio output device 104 (e.g., a speaker), and an impulse response h(n) representing an echo path between the audio output device 104 and the audio input device 102, where n is a frame index in the time domain. Acoustic echo cancellation system 100 estimates the acoustic echo signal ^^(n) by estimating the impulse response ℎ^(n) of the audio input device 102 and the audio output device 104 (e.g., the echo path). This echo is the output of the near- end signal x(n) (e.g., a reference signal) through the room impulse response h as provided by Equation 1: ^^(^) = ^^(^)^ [Equation 1]where: ^= ^ℎ^ ℎ^ … ℎ^^^^^ [Equation 2]^ = ^^(^) ^(^ − 1) … ^(^ − ^ + 1)^^ [Equation 3]X(n) is the length-L history of the far-end signal x(n) and L is the length of the echo path. The estimated echo signal ^^(n) is compared to the original microphone signal d(n) to generate an error signal e(n). For example, error signal e(n) may be determined as a difference (e.g., subtraction) between the original microphone signal d(n) and the estimated echo signal ^^(n).

[0023] The example echo cancellation system 100 includes adaptive filters (such as the main filter 204 and the shadow filter 206 described below with respect to FIG. 2) that may be adapted faster when initial convergence or echo path changes (for example, device movement changes, mute / unmute changes, volume changes, rendering angle changes, and the like) to model a new echo path. Alternatively, adaptive filters may be adapted more slowly (or even stopped altogether) during a double talk time state. Accordingly, embodiments described herein provide for a state detector (see FIG. 4) for use with acoustic echo cancellation. The state detector detects whether the echo cancellation system 100 is in a converging state, a double talk state, or a stable state. The state of the echo cancellation system 100 may impact the selection of echo re-filtering filters (see FIG. 4) and / or selected step sizes of adaptive filters. The use of a state detector to alter adaptive filters improves the performance of acoustic echo cancellation.

[0024] FIG. 2 illustrates a diagram of another example echo cancellation system 200. The example echo cancellation system 200 includes an audio input device 102 that provides original microphone signal d(n), an audio output device 104, and a controller 202. The echo cancellation system 200 is a dual-path filter system that includes two adaptive filters: a main filter (e.g., a first adaptive filter type) 204 and a shadow filter (e.g., a second adaptive filter type) 206. The main filter 204 may be a highly adaptive filter that determines filter coefficients responsive to current audio conditions (e.g., responsive to a current error signal, responsive to a detected change in the current audio conditions, etc.). The shadow filter 206 may be a conservative adaptive filter that provides little or no change in filter coefficients responsive to current audio conditions. The main filter 204 is adapted faster than the shadow filter 206 while the main filter 204 and the shadow filter 206 both copy each other. Further details regarding the main filter 204 and the shadow filter 206 can be found in PCT Application Publication No. WO 2022 / 120085, “SUBBAND DOMAIN ACOUSTIC ECHO CANCELLER BASED ACOUSTIC STATE ESTMIATOR,” which is herein incorporated by reference in its entirety. While the controller 202 is illustrated as being connected to the main filter 204 and the shadow filter 206, in some embodiments, the main filter 204 and the shadow filter 206 may be implemented by the controller 202. An example architecture that is suitable for use as controller 202 is described further below with respect to FIGS. 4A-4B.

[0025] The main filter 204 and the shadow filter 206 both estimate the impulse response ℎ^(n) ofthe audio input device 102 and the audio output device 104 to subsequently estimate the audio signal^^ (n). The main filter 204 detects a main filter impulse response ℎ^^(n) and outputs a main filterestimated echo signal ^^^(n). Similarly, the shadow filter 206 detects a shadow filter impulse response ℎ^^(n) and outputs a shadow filter estimated echo signal ^^^(n).

[0026] In various examples, main filter error ^^(n) and shadow filter error ^^(n) may be identified as provided below in Equations 4 and 5: ^^(^, ^) = (^, ^) − ^!^ (^, ^) [Equation 4]^^(^, ^) = (^, ^) − ^"^ (^, ^) [Equation 5]where l is the frame index and k is the frequency bin index.

[0027] The error difference between the main filter 204 and the shadow filter 206 may also be observed. For example, as provided by Equations 6-8: #$$%$ (^, ) | ^( )|)^ ^ = 10^%'10( ^ ^, ^ ) [Equation 6][Equation 7]∆#$$%$(^, ^) = #$$%$^(^, ^) − #$$%$^(^, ^) [Equation 8]

[0028] FIG. 3 is a plot illustrating the error difference ∆#$$%$ between a first adaptive filter type and a second adaptive filter type over a plurality of frames. For example, the plot illustrates the error difference ∆#$$%$ between an example main filter 204 and an example shadow filter 206 at each frame. The error difference ∆#$$%$ varies across each frame for different frequency bins. The error difference ∆#$$%$ indicates a state of the echo cancellation system 200. For example, a high negative error difference ∆#$$%$ before the 1000thframe (shown by 302) indicates a convergingstate (and, in particular, an initial convergence state). Similarly, a high negative error difference∆#$$%$ between the 5800th and 9000th frames (shown by 306) indicates a converging state (and, inparticular, an echo path changing state). A high positive error difference ∆#$$%$ between the 1000thand 4500thframes (shown by 304) indicate a double talk state. Remaining frames with low or near-zero error difference ∆#$$%$ indicate a stable state.

[0029] FIG. 4A shows a block diagram of an example electronic device architecture 450 suitable for implementing example embodiments of the present disclosure. Architecture 450 includes but is not limited to servers and client devices, systems, and methods as described in reference to systemFIGS. 1, 2, 4A, 5, 6, 7, and 8. As shown, the architecture 450 includes central processing unit (CPU) 451 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 452 or a program loaded from, for example, storage unit 458 to random access memory (RAM) 453. In RAM 453, the data required when CPU 451 performs the various processes is also stored, as required. CPU 451, ROPM 452, and RAM 453 are connected to one another via bus 454. Input / output (I / O) interface 455 is also connected to bus 454.

[0030] The following components are connected to I / O interface 455: input unit 456, that may include a keyboard, a mouse, or the like; output unit 457 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 458 including a hard disk, or another suitable storage device; and communication unit 459 including a network interface card such as a network card (e.g., wired or wireless).

[0031] In some implementations, input unit 456 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0032] In some implementations, output unit 457 includes systems with various number of speakers. Output unit 457 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0033] In some embodiments, communication unit 459 is configured to communicate with other devices (e.g., via a network). Drive 460 is also connected to I / O interface 455, as required. Removable medium 461, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 460, so that a computer program read therefrom is installed into storage unit 458, as required. A person skilled in the art would understand that although architecture 450 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0034] In accordance with example embodiments of the present disclosure, the processes described herein may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, thecomputer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 459, and / or installed from the removable medium 461, as shown in FIG. 4B.

[0035] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 451 in combination with other components of FIG. 4B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0036] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0037] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non- transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable ROM (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device,or any suitable combination of the foregoing.

[0038] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0039] FIG. 4B is a block diagram of an example of the controller 202. In the example illustrated, the controller 202 includes an electronic processor 400, a memory 402, and an input / output interface 404 connected over one or more control and / or data buses (for example, a communication bus 406). FIG. 4B illustrates only one example of the controller 202. The controller 202 may include more or fewer components and may perform functions other than those explicitly described herein.

[0040] In some examples, the electronic processor 400 is implemented as a microprocessor with separate memory, such as the memory 402. In other examples, the electronic processor 400 may be implemented as a microcontroller (with memory 402 on the same chip). In other examples, the electronic processor 400 may be implemented using multiple processors. In addition, the electronic processor 400 may be implemented partially or entirely as, for example, a field-programmable gate array (FPGA), an applications specific integrated circuit (ASIC), and the like and the memory 402 may not be needed or be modified accordingly. In the example illustrated, the memory 402 includes non-transitory, computer-readable memory that stores instructions that are received and executed by the electronic processor 400 to carry out the functionality of acoustic echo cancellation described herein. For example, the memory 402 may include instructions defining one or more of a state detector 408, an echo re-filter module 410, and / or a step size controller 412, each implemented by the electronic processor 400, as described below in more detail. The memory 402 may include, for example, a program storage area and a data storage area. The program storage area and the data storage area may include combinations of different types of memory, such as read-only memory and random-access memory.

[0041] The input / output interface 404 may include one or more input mechanisms, one or more output mechanisms, or a combination thereof. For example, the one or more input mechanisms may include the audio input device 102, and the one or more output mechanisms may include the audio output device 104.

[0042] The state detector 408 may be configured to detect the state of the echo cancellation system 200 based on the adaptive filter errors and also based on additional characteristics of the echo cancellation system 200, as will be described in further detail below with reference to FIGS. 5- 6. The echo re-filter module 410 may be configured to reapply the adaptive filters 604, as will be described in further detail below with reference to FIG. 6. The step size controller 412 may be configured to adjust a step size of the adaptive filters 604, as will be described in further detail below with reference to FIG. 6.

[0043] FIG. 5 is a block diagram illustrating an example state detection system 500. The example state detection system 500 includes an audio input device 102, a main filter 204, a shadow filter 206, a reference signal 506, and a state detector 408. The state detector 408, as illustrated, is configured to receive one or more signals as inputs from the audio input device 102, the main filter 204, the shadow filter 206, and the reference signal 506. For example, the inputs to the state detector 408 may include, but are not limited to, a main filter error signal Errorm received from the main filter 204 along a first input path 510, a shadow filter error signal Errors received from the shadow filter 206 along a second input path 512, a microphone power signal MicPower received from the audio input device along a third input path 514, and / or a reference power signal RefPower received along a fourth input path 516. The main filter error signal Errorm may be provided by mainfilter 204, which may indicate #$$%$^(^, ^) such as calculated in Equation 6 above. The shadowfilter error signal Errors may be provided by shadow filter 206, which may indicate #$$%$^(^, ^)such as calculated in Equation 7 above. In some instances, the controller 202 (e.g., see FIG. 4B)may be configured (e.g., by executable instructions) to calculate #$$%$^(^, ^) and #$$%$^(^, ^) andprovide the inputs to the state detector 408. The microphone power signal MicPower may be provided by the audio input device 102, and the reference power signal RefPower may be provided from reference signal 506 (e.g., x(n) in FIGS. 1-2). The MicPower may be defined, for example, by Equation 9: +,-.% / ^$(^, ^) = 10^%'10(| (^, ^)|)) [Equation 9]

[0044] In some examples, the state detector 408 may use the main filter error signal Errorm, the shadow filter error signal Errors, the microphone power signal MicPower, and the reference power signal RefPower to determine whether the echo cancellation system 200 is in a converging state, a double talk state, or a stable state. In the example of FIG. 5, the state detector 408 outputs a converging state signal along a first output path 518 when the echo cancellation system 200 is in a converging state, outputs a double talk state signal along a second output path 520 when the echo cancellation system 200 is in a double talk state, and outputs a stable state signal along a third output path 522 when the echo cancellation system 200 is in a stable state. However, fewer or more output paths may also be provided. For example, the state of the echo cancellation system 200 may be indicated by the state detector 408 along a single output path. The state detection result from the state detector 408 may be used to set or alter parameters associated with echo signal filtering and step size control of the adaptive filters (e.g., via control signals based on the state detection result). For example, if the state detector 408 detects that the echo cancellation system 200 is in the converging state, then echo signal filtering and step size control parameters of the adaptive filters are updated to accelerate convergence and remove additional echo. The example signals described above are merely illustrative examples, and additional or different signals may be employed by the state detection system 500 without departing from the spirit of the present disclosure.

[0045] FIG. 6 illustrates a block diagram representing an example acoustic echo cancellation algorithm 600. The acoustic echo cancellation algorithm 600 includes input signals 602, adaptive filters 604 (e.g., the main filter 204 and the shadow filter 206), the state detector 408, the echo re- filter module 410, and the step size controller 412. The input signals 602 may include the microphone power signal MicPower and the reference power signal RefPower. The input signals 602 are provided to both the adaptive filters 604 and the state detector 408 along a first path 606. The adaptive filters 604 provide the main filter error signal Errorm and the shadow filter error signal Errors to the state detector 408 along a second path 608, such as described previously with respect to FIG. 5. The state detector 408 provides control signals according to detected state information (e.g., Converging State Control Signals, Double Talk State Control Signals, etc.) to the echo re-filter module 410 and the step size controller 412. For example, Converging State Control Signals are provided to both the echo re-filter module 410 and the step-size controller 412 along a third path 610. The Double Talk State Control Signals are provided to the step size controller 412 along a fourth path 612. The outputs of the echo re-filter block 410 and the step size controller 412 are provided as a feedback signal along a fifth path 614 to the adaptive filters 604.

[0046] The state detector 408 is configured to identify one or more features that indicate different states of the echo cancellation system 200 based on one or more of the input signals 602 and / or one or more of the error signals Errormand Errors, responsively generate control signals, and transmit the generated control signals to one or more of the echo re-filter module 410 and / or the step size controller 412. When the state is a converging state, the control signals are adjusted by the state detector 408 such that the echo re-filter module 410 responsively reapplies the adaptive filters 604, and the step size controller 412 responsively updates the step size of the adaptive filters 604. When the state is a double talk state, the state detector 408 may be configured to only transmit control signals to the step size controller 412. The echo re-filter module 410 and the step size controller 412 responsively apply and update the adaptive filters 604 as indicated by the state detector 408.

[0047] Example features that are identified by the state detector 408 may include one or more of a main filter update counter, a frequency bin ratio of the main filter update counter, a shadow filter update counter, a frequency bin ratio of the shadow filter update counter, a positive frequency bin ratio of a smoothed delta error, and / or a negative frequency bin ratio of the smoothed delta error, to name a few.

[0048] To determine features, in some examples, the state detector 408 may first determine used frequency bins using the power of the reference signal: 0^1.% / ^$(^, ^) = 10^%'10(|^(^, ^)|)) [Equation 10]2 =×34567849 = 1 − exp (1 − ^?@? ) [Equation 11]Where N is the block The smoothed reference signal is provided by, for example, Equation 12: 0^1A.% / ^$ (^, ^) = 234567849∗ 0^1A.% / ^$ (^, ^ − 1) + C1 − 234567849D ∗ 0^1.% / ^$(^, ^) [Equation 12]

[0049] The order statistics version of the bins power is denoted as: EF^ G,^ = ^G,^H ^^, G,^H ^), … G,^H ^I^4JKL^=M^^ [Equation 13]Where: 0^1A.% / ^$ (G,^H ^^, ^) ≥ 0^1A.% / ^$ (G,^H ^), ^) ≥ ⋯ ≥ 0^1A.% / ^$ (G,^H ^I^4JKL^=M^, ^)[Equation 14] EF^ G,^PQR = S × .^$-^^T,^^^U [Equation 15]and K is a fixed value for a specific sampling rate. In this manner, calculations are focused on frequency bins where the energy is high, and frequency bins with lower energy may be ignored.

[0050] Next, the smoothed power of microphone signal d(k,l) is provided by Equation 16: +V-A.% / ^$ (EF^ G,^, ^) = 2WLX67849 ∗ +V-A.% / ^$ (EF^ G,^, ^ − 1) + (1 − 2WLX67849) ∗The smoothed main filter error for the used frequency bins is provided by Equation 17:#$A$%$^ (EF^ G,^, ^) = 2Y9979 ∗ #$A$%$^ (EF^ G,^, ^ − 1) + (1 − 2Y9979) ∗ #$A$%$^ (EF^ G,^, ^)The smoothed shadow filter error for the used frequency bins is provided by Equation 18: #A$$%$^ (EF^ G,^, ^) = 2Y9979 ∗ #A$$%$^ (EF^ G,^, ^ − 1) + (1 − 2Y9979) ∗ #A$$%$^ (EF^ G,^, ^)The smoothed error difference for the used frequency bins is provided by Equation 19:∆#A$$%$ (EF^ G,^, ^) = 2Y9979 ∗ ∆#A$$%$ (EF^ G,^, ^ − 1) + (1 − 2Y9979) ∗ ∆#A$$%$ (EF^ G,^, ^)

[0051] Echo return loss enhancement (ERLE) provides a value indicating how much echo was suppressed by a filter in the log domain. Using the smoothed microphone power and the smoothed error signals, the ERLE for the main filter and the shadow filter can be calculated using Equations 20 and 21: #0^#^(EF^ G,^, ^) = +V-A.% / ^$ (EF^ G,^, ^) − #$A$%$^ (EF^ G,^, ^) [Equation 20]#0^#^(EF^ G,^, ^) = +V-A.% / ^$ (EF^ G,^, ^) − #A$$%$^ (EF^ G,^, ^) [Equation 21]

[0052] The smoothed error signals, the smoothed microphone power and the ERLE values are used to determine the plurality of features. For example, the main filter update counter isdetermined according to: EZ [T^\%Q^T^$^(EF^ G,^, ^) =]EZ [T^\%Q^T^$^(EF^ G,^, ^ − 1) + 1 ,1 ∆#A$$%$ (EF^ G,^, ^) ≥ ^ℎ[ 2+[,^^U0, %Tℎ^$ / ,F^Where: Shad2MainThis a threshold associated with copy logic between the main filter 204 and the shadow filter 206.

[0053] Update ratio values for the main filter 204 and the shadow filter 206 may be determined to how many bins of the used bins are involved in the filter coefficient update. The frequency bin ratio of the main filter update counter is determined according to: EZ [T^0[T,% ( ^M^(I`Jab4c7M^b49d(I^4JKL^,e))^ ^) = I^4JKL^=M^ [Equation 23]

[0054] to:EZ [T^\%Q^T^$^(EF^ G,^, ^)EZ [T^\%Q^T^$^(EF^ G,^, ^ − 1) + 5 ,1 ∆#A$$%$ (EF^ G,^, ^) ≤ +[,^2^ℎ[ ^U [^ #0^#^(EF^ G,^, ^) ≤ #0^#^U^

[0055] The frequency bin ratio of the shadow filter update counter is determined according to: EZ [T^0[T,% ^M^(I`Jab4c7M^b49?(I^4JKL^,e)iI`Jab4c7M^b49jk)^(^) = [Equation 25]error to: ∆#$$%$.%F,T,l^0[T,%(^) = ^M^(∆YA9979 (I^4JKL^,e)i∆Y9979jk)[Equation 26]

[0057] The negative frequency bin ratio of smoothed delta error is determined according to: =^M^(∆YA9979 (I^4JKL^,e) m^∆Y9979jk)[Equation 27]

[0058] One or more of the plurality of features described above may be used by the statedetector 408 to determine the state of the echo cancellation system 200. Other features are also contemplated and considered within the spirit of the present disclosure.

[0059] FIG. 7 illustrates a block diagram of various example methods 700 for determining the state of the echo cancellation system 200. The methods 700 may include one or more of functional blocks, steps, operations, modules, or portions as illustrated by blocks 702, 704, 706, 708, 710, 712, and 714, which may be performed, for example, by the controller 202 implementing the state detector 408. The various process blocks illustrated in FIG. 7 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at decision block 702.

[0060] At decision block 702, the controller 202 determines whether the positive frequency bin ratio of smoothed delta error is greater than or equal to a ratio threshold RatioTh. When the positive frequency bin ratio of smoothed delta error is less than the ratio threshold RatioTh (“NO” at decision block 702), the controller 202 may proceed to decision block 704. When the positive frequency bin ratio of smoothed delta error is greater than or equal to the ratio threshold RatioTh(“YES” at decision block 702), the controller 202 may proceed to decision block 712.

[0061] At decision block 704, the controller 202 determines whether the negative frequency bin ratio of smoothed delta error is greater than or equal to the ratio threshold RatioTh. When the negative frequency bin ratio of smoothed delta error is less than the ratio threshold RatioTh(“NO” at decision block 704), the controller 202 may proceed to block 706 and determines that the state of the echo cancellation system 200 is a stable state. In some embodiments, when the state of the echo cancellation system 200 is the stable state, the current settings of the adaptive filters 604 are maintained. When the negative frequency bin ratio of smoothed delta error is greater than or equal to the ratio threshold RatioTh (“YES” at decision block 704), the controller 202 may proceed to decision block 708.

[0062] At decision block 708, the controller 202 determines whether the frequency bin ratio of the shadow filter update counter is greater than or equal to the ratio threshold RatioTh. When the frequency bin ratio of the shadow filter update counter is greater than or equal to the ratio threshold RatioTh(“YES” at decision block 708), the controller 202 may proceed to block 710 and determinesthat the state of the echo cancellation system 200 is a converging state. When the frequency bin ratio of the shadow filter update counter is less than the ratio threshold RatioTh (“NO” at decision block 708), the controller 202 may proceed to block 706 and determines that the state of the echo cancellation system 200 is the stable state.

[0063] At decision block 712, the controller 202 determines whether the frequency bin ratio of the main filter update counter is greater than or equal to the ratio threshold RatioTh. When the whether the frequency bin ratio of the shadow filter update counter is less than the ratio threshold RatioTh (“NO” at decision block 712), the controller 202 may proceed to block 706 and determine that the state of the echo cancellation system 200 is the stable state. When the whether the frequency bin ratio of the shadow filter update counter is greater than or equal to the ratio threshold RatioTh (“YES” at decision block 712), the controller 202 may proceed to block 714 and determine that the state of the echo cancellation system 200 is a double talk state. The double talk state may only exist in particular frequency bins, defined by Equations 28 and 29: ∆#A$$%$ (G,^H ^, ^) ≥ ∆#$$%$^U [Equation 28]EZ [T^\%Q^T^$^(G,^H ^, ^) ≥ EZ [T^\%Q^T^$^U [Equation 29]

[0064] In the illustrated example of FIG. 7, the ratio threshold RatioTh is the same at blocks 702, 704, 708, and 712. However, in other examples, the threshold at block 702, block 704, block 708, block 712, or a combination thereof may have a different value.

[0065] The methods 700 provide a particular heuristic rule-based example of determining the state of the echo cancellation system 200. In other instances, a machine learning model may be used to determine the state of the echo cancellation system 200. For example, the plurality of features (e.g., the main filter update counter, the frequency bin ratio of the main filter update counter, the shadow filter update counter, the frequency bin ratio of the shadow filter update counter, the positive frequency bin ratio of the smoothed delta error, and the negative frequency bin ratio of the smoothed delta error) are provided as inputs to a machine learning model implemented by the controller 202. The machine learning model then outputs the state of the echo cancellation system 200. The machine learning model may be, for example, a deep neural network (DNN), a convolutional neural network (CNN), an ensemble learning algorithm (for example an Adaptive Boosting algorithm), or the like.

[0066] When the state of the echo cancellation system 200 is the converging state, one or more of the plurality of features may be used to responsively update the adaptive filters 604 at the echo re- filter module 410 and the step size controller 412. FIG. 8 illustrates a block diagram of an acoustic echo cancelling algorithm 800 with echo re-filtering. The algorithm 800 includes blocks 802, 804, 806, 808, 810, and 812, as well as state detector 408. For various examples, processing of the various blocks of algorithm 800, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 802.

[0067] At block 802, the main filter 204 and the shadow filter 206 receive a microphone signal (e.g., signal d(n)) and a reference signal (e.g., reference signal 506). The main filter error signal Errormand the shadow filter error signal Errorsare used to both (i) update the adaptive filters and (ii) determine the state of the echo cancellation system 200 using the state detector 408.Additionally, the main filter estimated echo signal ^^^(n) and the shadow filter estimated echo signal^^^(n) are output by the main filter 204 and the shadow filter 206, as previously described withrespect to FIG. 2.

[0068] At block 804, the main filter 204 and the shadow filter 206 update their respective filter coefficients based on the main filter error signal Errormand the shadow filter error signal Errors. At block 806, the filter coefficients of the main filter 204 and the shadow filter 206 are compared. Filter coefficients with relatively small errors replace filter coefficients with relatively larger errors, thereby updating the main filter 204 and the shadow filter 206 to achieve smaller estimated error. The updated main filter 204 and the updated shadow filter 206 are then provided to block 810. In some instances, the operations of blocks 804 and 806 are only performed when the update counters defined in Equation 22 and Equation 24 are satisfied (e.g., the counters are above the respective thresholds).

[0069] The state detector 408 determines a detection result, and when the state detector 408 detects that the echo cancellation system 200 is in a converging state (at decision block 808), one or more of the plurality of features are used to update the main filter 204 and the shadow filter 206 to accelerate the adaptation process and remove additional echo by re-filtering the microphone signal at block 810. Otherwise, the initial result of the main filter 204 and the shadow filter 206 is used as the output (at block 812). The output of the acoustic echo cancelling algorithm 800 (at block 812) may be, for example, outputting the filtered signal via a speaker, storing the filtered signal in memory, or the like.

[0070] In an instance where the state of the echo cancellation system 200 is a double talk state, the step size of the adaptation of the adaptive filters 604 is adjusted by the step size controller 412 such that the adaptive filters 604 are more quickly adapted. For example, in some instances, the step size controller 412 uses the ratio of smoothed power of the estimated echo and the smoothed estimated error to control the step size µ, provided by Equation 30: n(^, ^) = o^77bU4JYXU767849(p,e)o^77bU4JY997967849(p,e) [Equation 30]

[0071] When usingfor determining the step size, the step size µ can be set to a smaller value during double talk time and a larger value during other states, as shown in FIG. 9. For example, in one implementation, when the state is a converging state: nL^b,n(^, ^) = q o^77bU4JYXU767849(p,e) L5 c7^r49sL^s obab4tbU498L^4 [Equation 31]where nis a double talk state, frequency bins at which double talk exists may be used to refine the step size result, as shown by Equation 32: u^,'ℎT o^77bU4JYXU767849(p,e)v^ ∗ o^77bU4JY997967849(p,e) ,[Equation 31]a

[0072] FIG. 10 illustrates a block diagram of various example methods 1000 for controlling the echo cancellation system 200. The methods 1000 may include one or more of functional blocks, steps, operations, modules, or portions as illustrated by blocks 1002, 1004, 1006, and 1008, which may be performed, for example, by the controller 202. The various process blocks illustrated in FIG. 10 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1002.

[0073] At block 1002, the controller 202 initiates or enables the system to perform a firstfiltering operation on an audio input signal received from an audio input device 102. For example, the adaptive filters 604 are applied to the audio input signal. Processing may then proceed from block 1002 to block 1004.

[0074] At block 1004, the controller 202 determines, based on one or more of a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of the acoustic echo cancellation system 100. For example, the controller 202 receives one or more of a plurality of features associated with the main filter 204 and / or the shadow filter 206. In some embodiments, one or more of the plurality of features are evaluated by the state detector 408 to determine the state of the system. The state of the acoustic echo cancellation system 100 may be, for example, a converging state, a double talk state, a stable state, or the like. Processing may then proceed from block 1004 to block 1006.

[0075] At block 1006, the controller 202 initiates or enables the system to perform, when in the converging state, a second filtering operation on the audio input signal. For example, the controller 202 determines that the state of the echo cancellation system 100 is the converging state. In response, the controller 202 enables (e.g., via control signals, communication messages, etc.) the echo re-filter module 410 to apply the adaptive filters 604 on the audio input signal a second, subsequent time. Processing may then proceed from block 1006 to block 1008.

[0076] At block 1008, the controller 202 adjusts (e.g. via control signals, communication messages, etc.), when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter. For example, when the controller 202 determines that the state of the echo cancellation system 100 is the double talk state, the controller 202 responsively initiates or enables the step size controller 412 to adjust a step size of the main filter 204, the shadow filter 206, or a combination thereof.

[0077] Systems, methods, and devices in accordance with the present disclosure may take any one or more of the following configurations.

[0078] (1) An audio processing method, comprising: performing a first filtering operation on an audio input signal received from an audio input device; determining, based on a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a convergingstate and a double talk state; performing, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal; and adjusting, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

[0079] (2) The audio processing method according to (1), further comprising: estimating a first echo signal using the first adaptive filter; and estimating a second echo signal using the second adaptive filter.

[0080] (3) The audio processing method according to (2), further comprising: determining an error of the first adaptive filter by subtracting the estimated first echo signal from the audio input signal; and determining an error of the second adaptive filter by subtracting the estimated second echo signal from the audio input signal.

[0081] (4) The audio processing method according to (3), wherein determining the state of the acoustic echo cancellation system is based on the error of the first adaptive filter and the error of the second adaptive filter.

[0082] (5) The audio processing method according to (4), further comprising: determining the plurality of features based on, at least one of frequency bins with high energy, a power of the audio input signal, the error of the first adaptive filter, and the error of the second adaptive filter.

[0083] (6) The audio processing method according to any one of (1) to (5), further comprising: determining the plurality of features based on an echo return loss enhancement of the first adaptive filter and an echo return loss enhancement of the second adaptive filter.

[0084] (7) The audio processing method according to any one of (1) to (6), wherein the state of the acoustic echo cancellation system is determined by a machine learning model.

[0085] (8) The audio processing method according to any one of (1) to (7), wherein determining a state of an acoustic echo cancellation system includes: identifying frequency bins in which the double talk state exists; and determining the state of the acoustic echo cancelling system based on the frequency bins in which the double talk state exists.

[0086] (9) The audio processing method according to any one of (1) to (8), further comprising:outputting the filtered audio input signal.

[0087] (10) The audio processing method according to (9), wherein outputting the filtered audio signal further comprises one of: storing the filtered audio input signal in a memory; transmitting the filtered audio input signal to a device; or providing the filtered audio input signal to a speaker.

[0088] (11) The audio processing method according to any one of (1) to (10), wherein the converging state indicates a negative error difference between the first adaptive filter and the second adaptive filter, and wherein the double talk state indicates a positive error difference between the first adaptive filter and the second adaptive filter.

[0089] (12) The audio processing method according to any one of (1) to (11), wherein the group of states includes a stable state, the method further comprising: maintaining the step size of the first adaptive filter and the step size of the second adaptive filter; and performing the first filtering operation on a second audio input signal received from the audio input device.

[0090] (13) A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method according to any one of (1) to (12).

[0091] (14) An apparatus for performing acoustic echo cancellation, the apparatus comprising: an input device configured to receive an audio input signal; and an electronic processor connected to the input device, the electronic processor configured to: perform a first filtering operation on the audio input signal received from the audio input device; determine, based on a plurality of features associated with a first adaptive filter and a second adaptive filter, a state of an acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state; perform, when in the converging state, a second filtering operation on the audio input signal to generate a filtered audio input signal; and adjust, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

[0092] (15) The apparatus according to (14), wherein the electronic processor is configured to: estimate a first echo signal using the first adaptive filter; estimate a second echo signal using the second adaptive filter; determine an error of the first adaptive filter by subtracting the estimated first echo signal from the audio input signal; and determine an error of the second adaptive filter bysubtracting the estimated second echo signal from the audio input signal, wherein the state of the acoustic echo cancellation system is determined based on the error of the first adaptive filter and the error of the second adaptive filter.

[0093] (16) The apparatus according to (15), wherein the plurality of features is determined based on, at least one of frequency bins with high energy, a power of the audio input signal, the error of the first adaptive filter, and the error of the second adaptive filter.

[0094] (17) The apparatus according to any one of (14) to (16), wherein the plurality of features is determined based on an echo return loss enhancement of the first adaptive filter and an echo return loss enhancement of the second adaptive filter.

[0095] (18) The apparatus according to any one of (14) to (17), wherein the state of the acoustic echo cancellation system is determined by a machine learning model implemented by the electronic processor.

[0096] (19) The apparatus according to any one of (14) to (18), wherein the converging state indicates a negative error difference between the first adaptive filter and the second adaptive filter, and wherein the double talk state indicates a positive error difference between the first adaptive filter and the second adaptive filter.

[0097] (20) The apparatus according to any one of (14) to (19), wherein the group of states includes a stable state, and wherein the electronic processor is configured to: maintain the step size of the first adaptive filter and the step size of the second adaptive filter; and perform the first filtering operation on a second audio input signal received from the audio input device.

[0098] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0099] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0100] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0101] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

[0102] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0103] Aspects of the systems described herein may be implemented in an appropriate computer-based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on variousdifferent network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0104] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non- transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0105] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.

[0106] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0107] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0108] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. Thesame applies to the term “implementation.”

[0109] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0110] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0111] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0112] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.

[0113] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitlyinclude, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0114] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0115] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0116] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0117] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

Claims

CLAIMS What is claimed is:

1. An audio processing method performed in an acoustic echo cancellation system, the method comprising: performing a first filtering operation on an audio input signal received from an audio input device using a first adaptive filter and a second adaptive filter; determining, based on at least one of the audio input signal and an output of the first filtering operation comprising a plurality of features associated with the first adaptive filter and the second adaptive filter, a state of the acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state; responsively generate a control signal based on the state of the acoustic echo cancellation system;performing, when in the converging state, a second filtering operation on the audio input signal by responsively reapplying the first and second adaptive filters based on the control signal to generate a filtered audio input signal, wherein a step size of at least one of the first adaptive filter and the second adaptive filter is adjusted based on the control signal; and adjusting, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter based on the control signal such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

2. The audio processing method of claim 1, further comprising: estimating a first echo signal using the first adaptive filter; and estimating a second echo signal using the second adaptive filter.

3. The audio processing method of claim 2, further comprising: determining an error of the first adaptive filter by subtracting the estimated first echo signal from the audio input signal; and determining an error of the second adaptive filter by subtracting the estimated second echo signal from the audio input signal.

4. The audio processing method of claim 3, wherein determining the state of the acoustic echocancellation system is based on the error of the first adaptive filter and the error of the second adaptive filter.

5. The audio processing method of claim 4, further comprising: determining the plurality of features based on, at least one of frequency bins with high energy, a power of the audio input signal, the error of the first adaptive filter, and the error of the second adaptive filter.

6. The audio processing method of any one of claims 1 to 5, further comprising: determining the plurality of features based on an echo return loss enhancement of the first adaptive filter and an echo return loss enhancement of the second adaptive filter.

7. The audio processing method of any one of claims 1 to 6, wherein the state of the acoustic echo cancellation system is determined by a machine learning model.

8. The audio processing method of any one of claims 1 to 7, wherein determining a state of an acoustic echo cancellation system includes: identifying frequency bins in which the double talk state exists; and determining the state of the acoustic echo cancelling system based on the frequency bins in which the double talk state exists.

9. The audio processing method of any one of claims 1 to 8, further comprising: outputting the filtered audio input signal.

10. The audio processing method of claim 9, wherein outputting the filtered audio signal further comprises one of: storing the filtered audio input signal in a memory; transmitting the filtered audio input signal to a device; or providing the filtered audio input signal to a speaker.

11. The audio processing method of any one of claims 1 to 10, wherein the converging state indicates a negative error difference between the first adaptive filter and the second adaptive filter, and wherein the double talk state indicates a positive error difference between the first adaptive filterand the second adaptive filter.

12. The audio processing method of any one of claims 1 to 11, wherein the group of states includes a stable state, the method further comprising: maintaining the step size of the first adaptive filter and the step size of the second adaptive filter; and performing the first filtering operation on a second audio input signal received from the audio input device.

13. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1 to 12.

14. An acoustic echo cancellation system comprising: an input device configured to receive an audio input signal; and an electronic processor connected to the input device, the electronic processor configured to: perform a first filtering operation on the audio input signal received from the audio input device using a first filter and a second adaptive filter; determine, based on at least one of the audio input signal and an output of the first filtering operation comprising a plurality of features associated with the first adaptive filter and the second adaptive filter, a state of the acoustic echo cancellation system, wherein the state is one selected from a group of states including a converging state and a double talk state; responsively generate a control signal based on the state of the acoustic echo cancellation system; perform, when in the converging state, a second filtering operation on the audio input signal by responsively reapplying the first and second adaptive filters to generate a filtered audio input signal, wherein a step size of at least one of the first adaptive filter and the second adaptive filter is adjusted based on the control signal; and adjust, when in the double talk state, a step size of at least one of the first adaptive filter and the second adaptive filter based on the control signal such that an adaptation rate of the at least one of the first adaptive filter and the second adaptive filter is altered.

15. The apparatus of claim 14, wherein the electronic processor is configured to: estimate a first echo signal using the first adaptive filter; estimate a second echo signal using the second adaptive filter; determine an error of the first adaptive filter by subtracting the estimated first echo signal from the audio input signal; and determine an error of the second adaptive filter by subtracting the estimated second echo signal from the audio input signal, wherein the state of the acoustic echo cancellation system is determined based on the error of the first adaptive filter and the error of the second adaptive filter.

16. The apparatus of claim 15, wherein the plurality of features is determined based on, at least one of frequency bins with high energy, a power of the audio input signal, the error of the first adaptive filter, and the error of the second adaptive filter.

17. The apparatus of any one of claims 14 to 16, wherein the plurality of features is determined based on an echo return loss enhancement of the first adaptive filter and an echo return loss enhancement of the second adaptive filter.

18. The apparatus of any one of claims 14 to 17, wherein the state of the acoustic echo cancellation system is determined by a machine learning model implemented by the electronic processor.

19. The apparatus of any one of claims 14 to 18, wherein the converging state indicates a negative error difference between the first adaptive filter and the second adaptive filter, and wherein the double talk state indicates a positive error difference between the first adaptive filter and the second adaptive filter.

20. The apparatus of any one of claims 14 to 19, wherein the group of states includes a stable state, and wherein the electronic processor is configured to: maintain the step size of the first adaptive filter and the step size of the second adaptivefilter; and perform the first filtering operation on a second audio input signal received from the audio input device.

Citation Information

Patent Citations

  • Subband domain acoustic echo canceller based acoustic state estimator

    WO2022120085A1

  • Echo cancellation system and method

    CN111277718A