Echo cancellation
By employing spatial analysis to determine echo probability, the patent addresses the inaccuracy of conventional echo detection, enhancing echo suppression in audio signals through improved adaptation rates and gains in acoustic echo cancellation systems.
Patent Information
- Application Number
- PCT/US2025/034351
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-14
- Filing Date
- 2025-06-19
- Publication Date
- 2025-12-26
AI Technical Summary
Conventional echo management systems inaccurately detect echo due to speaker non-linearities, leading to insufficient suppression of echo in audio signals.
Echo detection and suppression techniques utilizing spatial characteristics of audio signals, including power vectors and beamforming, to determine echo probability and adjust adaptation rates and suppression gains in acoustic echo cancellation systems.
Accurately detects echo conditions, improving signal-to-echo ratio by effectively suppressing echo and residual echo, even in the presence of speaker distortions.
Smart Images

Figure IMGF000010_0001 
Figure IMGF000021_0001 
Figure IMGF000024_0001
Abstract
Description
D23133EP ECHO CANCELLATION CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from PCT Application No. PCT / CN2024 / 100179 filed on 19 June 2024, and U.S. Provisional Application No. 63 / 680,685 filed on 8 August 2024, and EP 242063223.9 filed 14 October 2024, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] This disclosure pertains to systems, methods, and media for echo cancellation. BACKGROUND
[0003] In audio content, presence of echo may be annoying and distracting to listeners. Accordingly, it is useful to suppress echo in audio signals when detected. However, it can be difficult to accurately detect echo conditions in order to accurately suppress echo and not over suppress desired audio content. NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implementsD23133EP a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X − M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.D23133EP SUMMARY
[0008] Techniques for performing echo cancellation are provided herein.
[0009] According to some embodiments, a method for suppressing echoes in audio signals may involve receiving a reference signal indicative of an echo signal associated with an audio capture and playback device. The method may further involve receiving an audio signal from the audio capture and playback device. The method may further involve determining a probability of echo present in the audio signal based at least in part on a comparison of spatial characteristics of the reference signal and spatial characteristics of the audio signal. The method may further involve performing echo suppression on the audio signal based at least in part on the probability of echo present in the audio signal.
[0010] In some examples, the method may further involve: determining a power vector associated with the reference signal, the power vector associated with the reference signal being indicative of a direction of arrival of the reference signal; and determining a power vector associated with the audio signal, the power vector associated with the audio signal being indicative of a direction of arrival of the audio signal, wherein determining the probability of echo comprises comparing the power vector associated with the reference signal and the power vector associated with the audio signal. In some examples, comparing the power vector associated with the reference signal and the power vector associated with the audio signal comprises determining a cosine similarity.
[0011] In some examples, determining the probability of echo comprises applying a beamforming technique to the audio signal. In some examples, applying the beamforming technique may involve: generating a first beamformed signal by suppressing ambient sound in the audio signal while preserving echo signal in the audio signal; generating a second beamformed signal by suppressing the echo signal in the audio signal while preserving the ambient sound in the audio signal; determining a first power associated with the first beamformed signal and a second power associated with the second beamformed signal; and determining the probability of echo based on the first power and the second power.
[0012] In some examples, performing echo suppression on the audio signal comprises modifying an adaptation rate of an acoustic echo cancellation (AEC) system based on the probability of echo present in the audio signal. In some examples, the adaptation rate is proportional to the probability of echo present in the audio signal.D23133EP
[0013] In some examples, performing echo suppression on the audio signal comprises suppressing residual echo remaining in the audio signal after the audio signal has been processed by an adaptive echo cancellation (AEC) system. In some examples, suppressing residual echo comprises determining a suppression gain based at least in part on the probability of echo present in the audio signal.
[0014] In some examples, the method may further involve augmenting the audio signal after performing echo suppression to compensate for over-suppression of echo in the audio signal. In some examples, augmenting the audio signal comprises selecting an infill signal from one or more of: a microphone channel selected as least affected by echo; an output of a beamforming technique; or signal from frequency bands of the audio signal selected as least affected by echo. In some examples, the method may further involve scaling the infill signal based at least in part on a probability of residual echo after performing the echo suppression and a level of content in the audio signal. In some examples, the level of content in the audio signal is determined based on a smoothing factor that considers the probability of residual echo and a cross-correlation between the reference signal and the audio signal after performing the echo suppression.
[0015] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.
[0016] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0017] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features,D23133EP aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figures 1A, 1B, and 1C are diagrams illustrating an example audio capture and playback device in accordance with some embodiments.
[0019] Figure 1D illustrates a block diagram of an example echo cancellation system in accordance with some embodiments.
[0020] Figure 2A illustrates an example implementation of a system that utilizes spatial analysis for echo cancellation in accordance with some embodiments.
[0021] Figure 2B illustrates an example implementation of a system that utilizes echo probability to adjust gains of a residual echo suppression (RES) block in accordance with some embodiments.
[0022] Figure 2C illustrates an example implementation of a system that utilizes echo probably to adjust adaptation rate of an adaptive echo cancellation (AEC) system and gains of an RES block in accordance with some embodiments.
[0023] Figure 3A illustrates an example implementation of a system for determining an echo reference power vector in accordance with some embodiments.
[0024] Figures 3B illustrates an example implementation of a system for determining an echo probability using a power vector in accordance with some embodiments.
[0025] Figure 3C illustrates an example implementation of a system for utilizing a trained neural network to determine an echo probability in accordance with some embodiments.
[0026] Figure 4 illustrates an example implementation of a system for determining a probability of a residual echo in accordance with some embodiments.
[0027] Figure 5 illustrates an example implementation of a system for determining an echo probability using beamforming techniques in accordance with some embodiments.
[0028] Figure 6 illustrates an example implementation of a system for controlling an adaptation rate of an acoustic echo cancellation (AEC) system in accordance with some embodiments.D23133EP
[0029] Figures 7A illustrates an example implementation of a system for performing residual echo suppression in accordance with some embodiments.
[0030] Figure 7B illustrates an example implementation of a system for determining residual suppression gains using control logic in accordance with some embodiments.
[0031] Figure 8 is a flowchart of an example process for performing echo suppression in accordance with some embodiments.
[0032] Figure 9A shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0033] Figure 9B illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.
[0034] Figure 9C illustrates a schematic block diagram of an example CPU implemented in the device architecture of Figure 9B that may be used to implement various aspects of the present disclosure.
[0035] Like reference numbers and designations in the various drawings indicate like elements.D23133EP DETAILED DESCRIPTION OF EMBODIMENTS
[0036] Conventional echo management systems and echo suppression techniques typically detect echo by determining a correlation between output sent to device speakers with signal received by the device microphone(s), with a high correlation corresponding to a high echo probability. However, conventional techniques do not account for speaker non-linearities, which makes echo detection using conventional techniques inaccurate. For example, due to speaker distortions, which may be highly non-linear, there may be echo present in the signal received by the microphone(s) which is not detected in the correlation process. Because the echo is not detected due to speaker non-linearities, the echo cannot be sufficiently suppressed.
[0037] Disclosed herein are techniques for detecting echo based on spatial characteristics. In particular, the spatial characteristics of a reference signal representing an echo signal as observed by device microphones compared to the spatial characteristics of an audio signal received at the device microphones are used to determine an echo probability. The spatial characteristics of the reference signal may indicate an expected direction of arrival of the echo source, and the spatial characteristics of the audio signal may indicate a direction of arrival of the audio signal for which echo suppression is to be performed. In other words, the echo probability may be determined based on the degree to which the direction of arrival of the audio signal matches a known direction of arrival associated with an echo. Note that the spatial characteristics of the echo signal are device dependent and may encapsulate the particular relative locations of the microphone(s) and the speaker(s) of a particular audio capture and playback device (e.g., a particular mobile phone, laptop computer, tablet computer, etc.). Because speaker distortions and non-linearities do not change the spatial characteristics of the echo signal, the techniques disclosed herein may more accurately detect echo conditions compared to conventional techniques. Note that, while conventional techniques may solely rely on correlation to detect echo, which can lead to inaccurate echo suppression, as described above, the techniques disclosed herein which utilize spatial characteristics of the echo signal may be used in conjunction with techniques that utilize correlation of signals to detect echo.
[0038] The echo probability may be used to perform echo suppression. In some embodiments, echo probability may be used to control an adaptation rate of an acoustic echo cancellation (AEC) system. It may be desirable to utilize a faster adaptation rate in instances in which the echo probability is higher, and a lower adaptation rate with a lower echo probability. In some embodiments, echo probability may be used to determine gains for a residual echo suppression scheme in which residual echo after initial echo suppression by an AEC system is furtherD23133EP suppressed. Note that, in some embodiments, echo probability may be used to both perform initial echo suppression by setting adaptation rates of an AEC system and perform residual echo suppression by setting residual echo suppression gains based on the echo probability.
[0039] The techniques disclosed herein may effectively set the adaptation rate to be low or even 0 in instances in which “double talk” is detected in which there is both echo signal and desired signal (which may include noise). Moreover, the techniques disclosed herein may improve the signal to echo ratio (SER) by more effectively detecting echo compared to conventional techniques and setting adaptation rates and / or suppression gains based on a more accurate echo detection.
[0040] The techniques described herein may be implemented by one or more processors of an audio capture and playback device. Figure 1A is a diagram illustrating an example audio capture and playback device 100 in accordance with some embodiments. In the example shown in Figure 1A, audio capture device 100 is a mobile phone. Audio capture device 100 has a front side 102 and a back side 104. Front side 102 of mobile phone 100 is associated with two microphones, top microphone 106 and bottom microphone 108. Back side 104 of mobile phone 100 includes a third microphone, back microphone 110. Audio capture and playback device 100 includes a top microphone 106. Audio capture and playback device 100 also includes a speaker 112.
[0041] Note that each microphone may be configured to capture audio content from a particular direction associated with the location of the microphone with respect to mobile phone 100. Each speaker may be configured to output audio content. Note that echo may originate when audio content output by the speaker is received by the microphones of the device.
[0042] Figure 1B illustrates a side view of mobile phone 100 that includes top microphone 106.
[0043] Figure 1C illustrates a side view of mobile phone 100 that includes bottom microphone 108.
[0044] Figure 1D illustrates a block diagram of an example echo cancellation system 150. Echo cancellation system 100 includes a speaker 152, a microphone 154, an adaptive echo cancellation system (AEC) 156, and a summer block 158. The echo cancellation system 150 includes an input x(n) signal, which represents an input signal sent to speaker 152, and an enhanced signal e(n). The input signal x(n) is provided to speaker 152 and AEC 156. An estimated impulse response h(n) represents an impulse response associated with the echo path between speaker 152 and microphone 154. Microphone 154 generates a microphone signal d(n). AEC 156 generates an output signal represented as ^^^^^which represents an estimated echo signal. Microphone signal d(n) and theD23133EP estimated echo signal ^^^^^ are provided to summer block 158. Summer block 158 generates the enhanced signal e(n) by subtracting the estimated echo signal ^^^^^ from microphone signal d(n).
[0045] AEC 156 reduces the acoustic signal (generally represented herein as y(n)) by estimating the impulse response (generally represented herein as h(n)) of the loudspeaker-microphone system (e.g., the echo path). As illustrated in Figure 1D, the echo is the output of the far-end signal(represented as x(n)) through the room impulse response h as modeled by:^^^^ = ^^^^^ℎ
[0046] In the equation given above, h may be represented as ℎ = [ℎ ^^, ℎ^, … ℎ^^^] , and where^ = [^^^^^^^ − 1^ … ^^^ − ^ + 1^]^ . X(n) represents the history of length L of the far-endsignal x(n), and L is the length of the echo path.
[0047] The estimated echo signal, represented as ^^^^^ may be generated by an adaptive filter. The estimated echo signal may be a linear combination of several inputs at time index n, and may be determined by: ^^^^^ = ^^^^^^^^^
[0048] In the equation given above, W(n) represents the weight vector of the adaptive filter (e.g.,^^^^ = [^^^^^ ^^^^^ … ^ ^^^^^] ).
[0049] The weight vector ^^^^ may be updated using a Least Mean Squares (LMS) algorithm.For example, an error signal ^^^^^^ may be determined by subtracting the estimated echo signal^^^^^ from the echo signal ^^^^:^^^^^^ = ^^^^ − ^^^^^The weight vector of the adaptive filter ^^^^ may then be updated based on the error signal^^^^^^ according to:^ ^ = ^^ ^^ ^ + 1 ^^ +^^^ ^^^ ^^^^^^ ∗ ^^^^where u is a step size may a or some
[0050] It should be noted that Figure 1D illustrates a single microphone and a single speaker, however, this is only for simplicity. The equations and signal paths described above may be extended to a device that utilizes one or more microphones and / or one or more speakers. Moreover, the techniques disclosed herein that utilize spatial characteristics of captured signals to detect andD23133EP suppress echo are utilized with audio capture devices that have at least two microphones in order to utilize spatial differences between the echo and the desired signal. Such a device may utilize any suitable number of speakers, including one, two, three, etc.
[0051] Conventional AEC systems assume that speakers are linear. However, because the effects of the speaker are non-linear (e.g., due to speaker distortions), conventional AEC techniques that rely on echo detection based on whether a microphone signal and speaker signal are correlated may not effectively suppress echo.
[0052] Disclosed herein are techniques for echo suppression that utilize spatial analysis to determine a probability of echo in the signal. The echo probability may be used to determine filter weights for an AEC system, as shown in and described below in connection with Figure 2A. Use of spatial analysis may accurately detect echo more accurately than conventional techniques that utilize correlation to detect echo. Moreover, in some embodiments, spatial analysis may be used to suppress residual echo remaining in the signal after the signal has been processed by an AEC system, as shown in and described below in connection with Figure 2B. Note that, in some embodiment, the spatial analysis techniques described herein to determine echo probability may be used to both determine adaptive filter weights of an AEC system and to suppress residual echo after the signal has been processed by the AEC system, as shown in and described below in connection with Figure 2C.
[0053] Figure 2A illustrates an example implementation of a system 200 that utilizes spatial analysis for echo cancellation in accordance with some embodiments. System 200 includes short time Fourier transform (STFT) blocks 206, a spatial analysis block 208, control logic 210, an AEC 212 which generates an AEC output signal 214, and an RES block 216 which generates an output signal 218. Two STFT blocks 206 may receive speaker signal 202 and microphone signal 204, respectively, to generate frequency domain representations of speaker signal 202 and microphone signal 204. The frequency domain representation of microphone signal 204 may be provided to spatial analysis block 208. The output of spatial analysis block 208 is provided to control logic 210. The output of control logic 210 is provided to AEC 212. AEC 212 also receives the frequency domain representation of speaker signal 202 and microphone signal 206 from the two STFT blocks 206. AEC 212 generates AEC output signal 214 by performing initial echo suppression, which is provided to RES block 216. RES block 216 generates output signal 218 by performing residual echo suppression.D23133EP
[0054] As illustrated, speaker signals 202 and microphone signals 204 may be obtained. Speaker signals 202 represent the signal that was played by the speakers of the audio capture and playback device (e.g., the signal sent to the speakers). Speaker signals 202 are represented as x(n) in Figure 1D. Microphone signals 204 may be audio signals obtained with one or more microphones of the audio capture and playback device. Microphone signals 204 represent the speaker signals 202 as captured by the microphones of the audio capture and playback device. Microphone signals 204 are represented as d(n) in Figure 1D. In some embodiments, speaker signals 202 and / or microphone signals 204 may be transformed to a frequency domain using short- time Fourier transform instance 206.
[0055] As illustrated, the frequency domain representation of the microphone signals 204 may be provided to spatial analysis block 208. Spatial analysis block 208 may be configured to determine a probability that an echo exists in microphone signal 204. In general, spatial analysis may comprise comparing spatial characteristics of the reference signal and spatial characteristics of the audio signal. As illustrated, spatial analysis block 208 may utilize power-vector based techniques (shown in and described in more detail below in connection with Figures 3B and 3C) or a beam-forming technique (shown in and described below in connection with Figure 5) to determine the probability of echo in microphone signal 204.
[0056] Based on the echo probability, control logic 210 may determine an adaptation rate used by AEC block 212. For example, in some embodiments, the adaptation rate may be proportional to the probability of echo. As a more particular example, the adaptation rate may be reduced when the probability of echo is relatively low, and conversely, the adaptation rate may be increased when the probability of echo is higher. Note that, some echo probability values may indicate a “double talk” condition in which there is both echo and desired signal present. Control logic 210 may determine an adaptation rate that is highest when the echo probability indicates a likelihood of mostly echo or only echo in microphone signals 204, a lower rate when the echo probability indicates a combination of echo and desired signal in microphone signals 204, and a lowest rate when the echo probability indicates no echo or mostly desired signal in microphone signal 204. Note that example techniques that may be implemented by control logic 210 are shown in and described below in connection with Figure 6.
[0057] Based on the output of control logic 210, AEC 212 may filter microphone signal 204 to generate AEC output signal 214. In some embodiments, despite the echo cancellation performed by AEC 212, there may be residual echo present in AEC output signal 214. Accordingly, residualD23133EP echo suppression (RES) block 216 may take, as input, AEC output signal 214, and may perform additional echo suppression to generate output signal 218.
[0058] In some embodiments, rather than adjusting an adaptation rate of an AEC system, the echo probability may be used to determine gains of residual echo suppression (RES) block. Figure 2B illustrates an example implementation of a system 220 that utilizes echo probability to adjust gains of an RES block. System 220 short time Fourier transform (STFT) blocks 206, an AEC 212 which generates an AEC output signal 214, a spatial analysis block 228, control logic 230, and an RES block 216 which generates an output signal 238. Two STFT blocks 206 may receive speaker signal 202 and microphone signal 204, respectively, to generate frequency domain representations of speaker signal 202 and microphone signal 204. The frequency domain representations of speaker signal 202 and microphone signal 204 generated by STFT blocks 206 are provided to AEC 212, which generates AEC output signal 214 by performing initial echo suppression. AEC output signal 214 is provided to spatial analysis block 228. The output of spatial analysis block 228 is provided to control logic 230. An output of control logic 230 is provided to RES block 216. RES block 216 also receives, as input, AEC output signal 214. Based on the output of control logic 230 and AEC output signal 214, REC block 216 performs residual echo suppression to generate output signal 238.
[0059] As illustrated, AEC output signal 214 may be provided to spatial analysis block 228. Similar to what is described above in connection with spatial analysis block 208 of Figure 2A, spatial analysis block 228 may be configured to determine a probability of echo in AEC output signal 214. Spatial analysis block 228 may determine the echo probability using power-vector based techniques or a beamforming technique. The echo probability determined by spatial analysis block 228 may then be provided to control logic 230, which may be configured to determine one or more gains to be used by RES block 216. Example techniques for determining the gains are shown in and described below in connection with Figures 7A and 7B. RES block 216 may then generate output signal 238.
[0060] In some embodiments, the echo probability may be used to adjust both the adaptation rate of the AEC system and the gains of an RES block. Figure 2C illustrates an example implementation of a system 250 that utilizes echo probability to adjust adaptation rate of an AEC system and gains of an RES block. System 250 includes short time Fourier transform (STFT) blocks 206, first spatial analysis block 208, first control logic 210, an AEC 212 which generates an AEC output signal 214, a second spatial analysis block 230, second control logic 228, and an RES block 216 which generates an output signal 238. Two STFT blocks 206 may receive speakerD23133EP signal 202 and microphone signal 204, respectively, to generate frequency domain representations of speaker signal 202 and microphone signal 204. The frequency domain representation of microphone signal 204 is provided to first spatial analysis block 208. The output of first spatial analysis block 28 is provided to first control logic 210. The output of first control logic 210 is provided to AEC 212. The frequency domain representations of speaker signal 202 and microphone signal 204 generated by STFT blocks 206 are also provided as input to AEC 212, which generates AEC output signal 214 by performing initial echo suppression based on the output of first control logic 210. AEC output signal 214 is provided to second spatial analysis block 230. The output of second spatial analysis block 230 is provided to control logic 228. An output of control logic 228 is provided to RES block 216. RES block 216 also receives, as input, AEC output signal 214. Based on the output of control logic 228 and AEC output signal 214, REC block 216 performs residual echo suppression to generate output signal 258.
[0061] As illustrated, first spatial analysis block 208 and first control logic 210 may determine an echo probability in microphone signal 204 and may determine an adaptation rate for AEC block 212 based on the echo probability. Second spatial analysis block 230 and second control logic 228 may determine an echo probability associated with AEC output signal 214 and may determine gains for RES block 216 based on the echo probability. The output of RES block 216 is output signal 258.
[0062] As described above, in some embodiments, echo probability may be determined based on power vectors. In particular, a power vector associated with a reference signal representative of an echo may be compared to a power vector associated with a microphone audio signal. Each power vector is indicative of a direction of arrival of the associated signal. Therefore, comparison of the power vector associated with the reference echo signal with the power vector associated with the microphone audio signal may indicate whether the direction of arrival of the microphone audio signal is similar to or the same as the expected direction of arrival of an echo signal. The echo probability may be determined based on the comparison of the power vectors such that a higher probability corresponds to microphone audio signal directions that are closer to the expected echo direction.
[0063] Note that the power vector associated with the echo reference signal represents the direction of arrival of an echo signal as observed by the microphones of the audio capture device. In other words, the power vector associated with the echo reference signal represents a characteristic of audio played from the speakers of the audio capture and playback device as observed by the microphones of the same device. In some embodiments, the power vectorD23133EP associated with the echo reference signal may be determined during a tuning stage or process that occurs prior to an inference stage. In some embodiments, the power vector may be determined using a three-dimensional model of the audio capture device that indicates relative distances between speakers and microphones, an acoustic simulation that identifies a transfer function associated with a particular device, and / or by using a real audio capture device and recording signals from microphones responsive to playing signals from the speakers.
[0064] Figure 3A illustrates an example implementation of a system 300 for determining an echo reference power vector in accordance with some embodiments. System 300 includes an STFT block 304, and a covariance matrix block 306, which generates an echo reference power vector 308. The set of n microphone signals 302 are provided as input to STFT block 304, which generates frequency domain representations of the n microphone signals 302. The output of STFT block 304 is provided to covariance matrix block 306, which generates, as an output, echo reference power vector 308.
[0065] Note that the technique illustrated in Figure 3A may be obtained during a tuning stage or process, and the obtained power vector may then be stored for later use at during an inference stage (e.g., to determine an echo probability associated with a microphone signal and to perform echo suppression based on the echo probability). As illustrated, n microphone signals 302 that capture an echo signal may be provided to STFT block 304. The frequency domain representation of the microphone signals 302 generated by STFT block 304 may be provided to covariance matrix block 306, which may generate echo reference power vector 308. Note that covariance matrix block 306 may be configured to determine covariance between the n microphone channels on a per-frequency band basis to determine echo reference power vector 308.
[0066] Figure 3B illustrates an example system 320 for determining an echo probability using a power vector in accordance with some embodiments. As illustrated, system 320 includes an STFT block 324, a covariance matrix block 326 which generates a power vector 328, and a cosine similarity block 330 that takes as input echo reference power vector 308 and power vector 328 and generates a probability that an echo exists. The set of n microphone signals 322 are provided as input to STFT block 324 which generates a frequency domain representation of the n microphone signals 322. The frequency domain representation is provided as input to covariance matrix block 326, which generates, as an output, power vector 328. Power vector 328 is provided as input to cosine similarity block 330, which additionally receives echo reference power vector 308 (e.g., as determined during a tuning stage, as shown in and described above in connection with Figure 3A), and generates, as an output, a probability that an echo exists in the n microphone signals 322.D23133EP
[0067] As illustrated, set of n microphone signals 322 may be obtained. Microphone signals 322 may represent audio signal captured by an audio playback and capture device during an inference stage, where the degree of echo in microphone signals 322 is unknown at the time of capture. Microphone signals 322 may be provided to STFT instance 324 which may be configured to generate a frequency domain representation of the n channels of microphone signals. The frequency domain representation may be used to generate power vector 328 using covariance matrix block 326. Similar to what is described above in connection with Figure 3A, covariance matrix block 326 may be configured to determine covariance between the n microphone channels on a per-frequency basis to determine power vector 328. Cosine similarity block 330 may be configured to compare echo reference power vector 308 (which may be determined using the techniques described above in connection with Figure 3A) and power vector 328 to generate the echo probability. As described above, the echo probability generated by cosine similarity block 330 corresponds to the probability that an echo exists in microphone signals 322.
[0068] In some embodiments, rather than comparing the echo reference power vector and the power vector associated with the microphone signals using cosine similarity, the echo reference power vector and the microphone signal power vector may be provided to a trained model (e.g., a trained neural network, a trained regression model, a trained classifier, etc.) configured to generate the echo probability as an output.
[0069] Figure 3C illustrates an example system 350 that utilizes a trained neural network to determine the echo probability in accordance with some embodiments. System 350 includes an STFT block 324, a covariance matrix block 326 which generates a power vector 328, and a trained neural network 352 that takes, as input, echo reference power vector 308 and power vector 328 and generates a probability that an echo exists. The n microphone signals 322 are provided as input to STFT block 324 which generates a frequency domain representation of the n microphone signals 322. The frequency domain representation is provided as input to covariance matrix block 326, which generates, as an output, power vector 328. Power vector 328 is provided as input to neural network 352, which additionally receives echo reference power vector 308 (e.g., as determined during a tuning stage, as shown in and described above in connection with Figure 3A), and generates, as an output, a probability that an echo exists in the n microphone signals 322.
[0070] Note that a trained model may be trained in any suitable manner. For example, the model may be trained using a training set that is composed of power vector training samples, each annotated with a ground truth annotation indicative of whether or not there is echo present in theD23133EP microphone signal associated with the power vector training sample. The model may then be trained to predict the corresponding ground truth label.
[0071] In some embodiments, the echo reference power vector may be compared to a power vector associated with an audio signal after processing by an AEC block to determine the probability of a residual echo in the audio signal after processing by the AEC block. In some embodiments, comparison of the echo reference power vector to the power vector of the post-AEC processed audio signal may be implemented using cosine similarity, or a trained machine learning model, as described above.
[0072] Figure 4 illustrates an example system 400 for determining a probability of residual echo in accordance with some embodiments. System 400 includes an STFT block 324, an AEC 212, a covariance matrix block 402 which generates a power vector 404, and a cosine similarity block 406 which takes power vector 404 and an echo reference power vector 308 as input and generates a probability of a residual echo. As illustrated, a set of n microphone signals 322 may be transformed to the frequency domain using STFT block 324. The frequency domain representation may be provided to AEC block 212 which may perform initial echo cancellation on the microphone signals. The output of AEC block 212 is provided to covariance matrix block 402. The output of AEC block 212 may comprise modified audio signals (e.g., after initial echo cancellation has been performed by AEC block 212), and may be used to determine power vector 404 using covariance matrix block 402. Cosine similarity block 406 receives, as input, echo reference power vector 308 and power vector 404. Cosine similarity block 406 generates, as output, a probability of residual echo by comparing power vector 404 with echo reference power vector 308 to determine the probability of residual echo in the modified audio signals after initial echo cancellation by AEC block 212.
[0073] In the example shown in Figure 4, power vector 404 and echo reference power vector 308 are compared using cosine similarity block 406. However, in some embodiments, cosine similarity block 406 may be replaced with a machine learning model that is trained to compare power vectors and generate, as an output, an echo probability based on the two power vectors.
[0074] Figures 3A, 3B, 3C, and 4 illustrate use of power vectors to determine the probability of an echo, whether the probability of an echo is with respect to the original microphone signals (prior to initial echo cancellation by an AEC block) or with respect to initially processed signals by an AEC block to determine the probability of residual echo. In some embodiments, rather than using a power vector based approach, echo probability may be determined using beamformingD23133EP techniques. Note that beamforming techniques may be used to determine the echo probability associated with the original microphone signals, or a residual echo probability after initial echo cancellation by an AEC block. Use of beamforming techniques to determine echo probability may involve generating a first beamformed signal by suppressing ambient sound in the audio signal while preserving echo signal in the audio signal, and generating a second beamformed signal by suppressing the echo signal in the audio signal while preserving the ambient sound in the audio signal. In other words, the first beamformed signal may represent mostly or entirely the echo portion while the second beamformed signal may represent mostly or entirely the ambient sound portion. A first power associated with the first beamformed signal and a second power associated with the second beamformed signal may be determined, and the echo probability may be determined based on the first power and the second power.
[0075] Figure 5 illustrates an example system 500 for determining echo probability associated with original microphone signals using beamforming techniques in accordance with some embodiments. System 500 includes an STFT block 324, a first beamforming block 502, a second beamforming block 504, a first power block 506, a second power block 508, a summer block 509, and a sigmoid function block 510. As illustrated, n microphone signals 322 may be transformed to the frequency domain using STFT block 324. First beamforming block 502 may generate a first beamformed signal that suppresses the ambient sound, and second beamforming block 504 may generate a second beamformed signal that suppresses the echo signal. The first beamformed signal may be provided as input to first power block 506, which determines a power associated with the first beamformed signal. The second beamformed signal may be provided as input to a second power block 508, which determines a power associated with the second power signal. The outputs of first power block 506 and second power block 508 are provided to summer block 509, which determines a difference between the power associated with the first beamformed signal and the power associated with the second beamformed signal. The output of summer block 509 is provided as input to sigmoid function block 510, which generates, as output, a probability that echo exists in the n microphone signals 322.
[0076] With reference to first beamforming block 502 and second beamforming block 504, both beams may be generated using the minimum variance distortionless response (MVDR) algorithm. In the MVDR algorithm, filter coefficients may be determined by: #^^ ∗ $D23133EP
[0077] In the equation given above, R represents the noise covariance, γ represents the steering vector, and (*)Hrepresents the conjugate transpose operator. The frequency domain representation of the microphone signals may be filtered using the MVDR weights.
[0078] For the first beamformed signal that suppresses the ambient sound, the relative transfer function (RTF) of the echo source may be used as the steering vector. The noise covariance may be estimated by iteratively smoothing the instant covariance, represented as Rinstant. In some examples, the noise covariance may be determined by: #= & ∗ # + ^1 − &^ ∗ #'()*+(*
[0079] In the equation given above, α represents a smoothing factor. The smoothing factor may be determined by the power of the reference signal. In one example, the smoothing factor may be determined by: &= 0 -. ^^.^^^^ / ^ 01^^^ ≥ 3ℎ^^4ℎ156, ^54^ & = 1
[0080] In the equation given above, the threshold may be, e.g., -70 dB, -60 dB, -50 dB, etc.
[0081] For the second beamformed signal that suppresses the echo, the steering vector may be determined by the microphone and speaker locations to minimize the impact from the echo signal. In one example, the steering vector may correspond to the forward looking direction, because the forward looking direction may have the lowest echo power. The noise covariance may be the echo covariance (e.g., used to calculate the echo reference power vector as described above in connection with Figure 3A). In some embodiments, the steering vector may be determined prior to inference time (e.g., as part of a tuning stage or a tuning process). In some embodiments, the noise covariance may be determined prior to inference time. Alternatively, in some embodiments, the noise covariance may be determined dynamically using: #= α ∗ # + ^1 − α^ ∗ #'()*+(*& = 0 -. ^^.^^^^ / ^ 01^^^ ≥ 3ℎ^^4ℎ156, ^54^ & = 1
[0082] Note that the threshold used to determine the smoothing factor for the second beamformed signal may be the same or different than the threshold used to determine the smoothing factor for the first beamformed signal.
[0083] After generating the first beamformed signal and the second beamformed signal, the power of each may be calculated using power blocks 506 and 508. Power may be determined inD23133EP decibels (dB). The sigmoid function block 510 may be used to limit the output between zero and one to determine the echo probability.
[0084] Note that the beamforming technique shown in and described above in connection with Figure 5 may be utilized to determine the residual echo probability instead of and / or in addition to the echo probability associated with the original signals. For example, a first beamformed signal and a second beamformed signal may be generated using the techniques described above in connection with Figure 5 using the output of an AEC block. Power may be determined for each of the first beamformed signal and the second beamformed signal, and the output of the sigmoid function may then correspond to the residual echo probability.
[0085] As described above, in some embodiments, the echo probability may be used to control an adaptation rate of an AEC block. In general, the adaptation rate may be proportional to the echo probability such that the adaptation rate is faster with relatively higher echo probability and lower (or 0) with relatively lower echo probability. Note that the echo probability may be determined using either a power vector based technique (e.g., as shown in and described above in connection with Figures 3A-3C) or a beamforming technique (e.g., as shown in and described above in connection with Figure 5). Additionally, it should be noted that the echo probability as initially determined may be a mono echo probability that indicates a likelihood of echo across all n microphone channels. The echo probability may then be expanded to n channels. Expansion of the echo probability may be based on the echo covariance matrix, e.g., as used to determine the echo reference power vector as shown in and described above in connection with Figure 3A. The expanded echo probability may be multiplied with a value of an update flag, which may indicate whether or not the adaptation rate is to be changed. The multiplied value may be used to control the adaptation rate of the AEC block.
[0086] Figure 6 depicts an example system 600 for controlling an adaptation rate of an AEC system in accordance with some embodiments. System 600 includes an STFT block 304, an AEC 212, an expansion block 602, and update flag block 604, and a multiplication block 606. As illustrated, n microphone signals 302 may be provided as input to STFT block 304, which transforms the microphone signals to the frequency domain. The output of STFT block 304 is provided to AEC 212. AEC 212 also receives, as input, an echo reference signal (which may be obtained and / or stored during a tuning stage or a setup stage for the audio capture and playback device). Expansion block 602 may receive, as input, a probability that echo exists in the n microphone signals 322. Update flag block 604 receives, as input, a reference power vector. The outputs of expansion block 602 and update flag block 604 are provided as input to multiplicationD23133EP block 606 to determine an adaptation rate for AEC 212. Multiplication block 606 is coupled to AEC 212 to control the adaptation rate of AEC 212.
[0087] With reference to expansion block 602, the echo probability (which may be determined using either power vector based techniques or a beamforming technique) may be provided to expansion block 602. Expansion block 602 may be configured to transform the mono echo probability to n channels corresponding to the n microphone signals. Expansion block 602 may multiply the echo probability by the echo power for the given microphone channel, where the echo power is determined from the corresponding element in the diagonal of the covariance matrix used to determine the echo reference power vector (e.g., covariance matrix 306 of Figure 3A). In one example, the echo probability for microphone channel i may be determined by: 0^ / ℎ10^1898-5-3^ = :1^1 '' ^ / ℎ10^1898-5-3^ ∗max ^0^, … 0>^
[0088] Inecho probability associated with the set of microphone signals as determined using either the power vector technique or the beamforming technique, as shown in Figures 3A-3C or Figure 5, respectively. Additionally, pirepresents the echo power of the ithmicrophone as determined from the covariance matrix.
[0089] With respect to update flag block 604, the echo reference power may be provided to update flag block 604. The update flag value may be updated based on the echo reference power, and may indicate whether or not the adaptation rate should be adapted. For example, the update flag may be set based on the echo reference power exceeding a threshold. Note that update flag may be a Boolean value indicating whether or not the adaptation rate should be adapted.
[0090] With respect to multiplication block 606, multiplication block 606 may be configured to multiply the update flag (which may be either 0 or 1) with the expanded echo probability to determine an adaptation rate for each of the n microphone channels. In one example, the echo adaptation rate may be determined by: ^ / ℎ19690393-1^ ^93^ = ^0693^ .59? ∗ ^ / ℎ10^1898-5-3^'
[0091] Note that the same adaptation rate may be applied to all microphone channels. The echo adaptation rate may then be provided to AEC block 212 to perform filtration of the frequency domain representation of the microphone signals 322.D23133EP
[0092] As shown in and described above in connection with Figure 2B, in some embodiments, a residual echo probability may be used to determine gains of a residual echo suppression (RES) block. The RES block may further suppress echo that was not fully suppressed by an AEC block.
[0093] Figure 7A illustrates an example system 700 for performing residual echo suppression. System 700 includes an AEC 212, an expansion block 702, a subtraction block 704, and a multiplication block 706. Expansion block 702 takes, as input, a probability of residual echo. The output of expansion block 702 represents the expanded residual echo probability for n channels. The output of expansion block 702 is provided as input to subtraction block 704. Subtraction block 704 subtracts the probability of residual echo from 1 to determine an output representing residual echo suppression gains. The output of subtraction block 704 is provided to multiplication block 706. Multiplication block 706 also receives, as input, an output of AEC 212. Multiplication block 706 applies the residual echo suppression gains to the output of AEC 212 to generate, as an output, an RES output, which corresponds to residual echo suppression on the output of AEC 212.
[0094] With reference to AEC 212 in Figure 7A, as illustrated, an output of an AEC 212 may be obtained, which corresponds to the microphone audio signal after initial echo suppression performed by AEC 212. A residual echo probability may be determined. The residual echo probability may be determined using power vector based techniques or using beamforming techniques. The residual echo probability may indicate a likelihood of echo remaining in the signal after processing by AEC 212. The residual echo probability may be expanded to n microphone channels using expansion block 702. Expansion block 702 may function similarity to expansion block 602 of Figure 6. For example, expansion block 702 may multiply the residual echo probability by the echo power for the given microphone channel, where the echo power is determined from the corresponding element in the diagonal of the covariance matrix used to determine the echo reference power vector (e.g., covariance matrix 306 of Figure 3A). The expanded residual echo probability may be subtracted at 704 to determine residual echo suppression gains, which may be applied via multiplication block 706 to the AEC output to generate a further suppressed RES output signal. In particular, expansion block 702 determines the ratio of echo to signal in each channel, and subtraction block 704 subtracts the echo ratio from one to determine how much signal should be preserved.
[0095] In some embodiments, the residual suppression gain may be determined from a correlation between the reference signal and the residual echo, as well as the residual echo probability.D23133EP
[0096] Figure 7B illustrates an example implementation of a system 750 for determining residual echo suppression gains using control logic in accordance with some embodiments. As illustrated, system 750 includes an AEC 212, an expansion block 702, a multiplication block 706, and control logic 752. Expansion block 702 takes, as input, a probability that a residual echo exists in the output of AEC 212. The output of expansion block 702 is provided to control logic 752. Control logic 752 also receives, as input, a correlation between the reference signal and the AEC output. Control logic 752 generates residual suppression gains, which are provided as input to multiplication block 706. Multiplication block 706 additionally receives, as input, an output of AEC 212. Multiplication block 706 generates an output signal RES output by applying the residual echo suppression gains to the output of AEC 212.
[0097] In one example, the suppression gains may be determined by: 4^00^^44-1^ ?9-^ = min^1 − ^B , 1 − ^C^
[0098] In the equation given above, xerepresents the correlation between the reference signal and the residual echo, and ep represents the residual echo probability. The suppression gains may then be applied to the output of AEC 212, similar to what is described above in connection with Figure 7A, to generate the RES output signal.
[0099] In some cases, echo suppression may over suppress the echo signal which can cause discontinuities and / or loss of desired portions of the audio signal. In some embodiments, the level of local content may be estimated and filled in with signal from another source. The other source may include the microphone channel identified as least affected by echo, the second beamformed signal (e.g., as shown in and described above in connection with Figure 5, which corresponds to the signal with echo suppressed), and / or surround frequency bands that are determined to be less affected by echo. In some embodiments, local content level may be estimated using a recursively smoothed periodogram P. By way of example, the periodogram may be determined by: D^., 3^ = & ∗ D^., 3 − 1^ + ^1 − &^ ∗ |E^., 3^|F
[0100] In the equation given above, |E^., 3^|F represents the instant periodogram. In theequation given above, α represents a smoothing factor. Example values of the smoothing factor include 0.85, 0.9, 0.95, 0.98, etc. The smoothing factor may be determined based on the residual echo probability and / or a cross-correlation between the reference signal and the output of the AEC that includes the residual echo. The cross correlation may be performed on a per-microphone basisD23133EP . In one example, the smoothing factor α is the maximum of the residual echo probability and the cross-correlation.
[0101] The infill signal that is used to fill in the signal may be generating by selecting a source, as described above. Generally the source may be one that is less affected by echo, and / or a beamformed signal that has suppressed echo. By way of example, the infilled signal may be generated by: G^.-55HI* = J ∗ #KLHI* + M1 − J N ∗ -^.-55'(CI*
[0102] In thesource used to infill the signal, and RESout represents the over-suppressed output of the RES. In the equation given above, β is a weighting parameter which may be determined based on the periodogram and the over- suppressed RES output. In some embodiments, beta may be the energy ration in the output infill signal, Infillout. In one example, β may be determined by: 1, #KLHI* > D^., 3^
[0103] Turning to Figure 8,process 800 for performing echo suppression is shown in accordance with some embodiments. Blocks of process 800 may be executed by one or more processors and / or control systems of a user device, such as an audio capture and playback device (e.g., a mobile phone, a tablet computer, a desktop computer, a laptop computer, etc.). An example of such a control system is control system 910 as shown in Figure 9A. In some embodiments, blocks of process 800 may be executed in an order other than what is shown in Figure 8. In some embodiments, two or more blocks of process 800 may be executed substantially in parallel. In some embodiments, one or more blocks of process 800 may be omitted.
[0104] Process 800 can begin at 802 by receiving a reference signal indicative of an echo signal associated with an audio capture and playback device. Note that, in some embodiments, the reference signal may be obtained and / or characterized during an initialization or tuning stage of the audio capture and playback device. For example, the reference signal may be obtained during a set up process and / or during the design process. In some embodiments, the reference signal may indicate characteristics of an echo when audio is played from speakers of the audio capture and playback device and captured by one or more microphones of the same device.D23133EP
[0105] At 804, process 800 can receive an audio signal from the audio capture and playback device. For example, the audio signal may be received by one or more microphones of the device. Note that block 804 may occur during an inference stage, and accordingly, the received audio signal may include any combination of echo and desired sound portions (e.g., all echo, no echo, or any degree of echo in between).
[0106] At 806, process 800 can determine a probability of echo present in the audio signal based at least in part on a comparison of spatial characteristics of the reference signal and spatial characteristics of the audio signal. In some embodiments, process 800 can determine the probability of echo present in the audio signal by comparing an estimated direction of arrival associated with a known echo signal (as indicated by the reference signal received or obtained at block 802) with the direction of arrival associated with the audio signal received at block 804. By way of example, in an instance in which the spatial characteristics of the reference signal align with the spatial characteristics of the audio signal such that it is likely the audio signal arrives from the direction expected of an echo signal, the echo probability may be correspondingly high.
[0107] As described above, the echo probability may be determined by comparing a power vector associated with the reference signal with a power vector associated with the audio signal (e.g., as described above in connection with Figures 3A-3C). Alternatively, the echo probability may be determined using beamforming techniques, as shown in and described above in connection with Figure 5. Additionally, it should be noted that the echo probability may be associated with the original microphone signals (e.g., as shown in Figure 2A), may be a residual echo probability indicating a probability of echo remaining after initial echo suppression by an AEC (e.g., as shown in Figure 2B), or may be two different probabilities comprising a first probability of echo associated with the original microphone signals and a residual echo probability indicating a probability of echo remaining after initial echo suppression (e.g., as shown in Figure 2C).
[0108] At 808, process 800 may perform echo suppression on the audio signal based at least in part on the probability of echo present in the audio signal. For example, in some embodiments, the echo probability may be used to determine an adaptation rate of an AEC block used to perform echo suppression. As another example, in some embodiments, the echo probability may be used to determine residual echo suppression gains to be applied after initial suppression using the AEC block. Note that, in some embodiments, echo probabilities may be used to perform both an initial suppression using the AEC block and residual echo suppression, although the echo probability used to perform residual echo suppression is determined using an output signal of the AEC block (e.g., as shown in and described above in connection with Figure 4).D23133EP
[0109] In some embodiments, after performing echo suppression, the resulting signal may be infilled to account for any over suppression of desired content during the echo suppression process.
[0110] After echo suppression (and optionally, a signal infill process to account for over suppression of desired content), the resulting audio signal may be played back, transmitted to another device (e.g., another user device, a server or cloud device, etc.), stored for later playback, or any combination thereof.
[0111] Figure 9A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 9A are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 900 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 900 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0112] According to some alternative implementations the apparatus 900 may be, or may include, a server. In some such examples, the apparatus 900 may be, or may include, an encoder. Accordingly, in some instances the apparatus 900 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 900 may be a device that is configured for use in “the cloud,” e.g., a server.
[0113] In this example, the apparatus 900 includes an interface system 905 and a control system 910. The interface system 905 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 905 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 900 is executing.
[0114] The interface system 905 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may includeD23133EP spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0115] The interface system 905 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 905 may include one or more wireless interfaces. The interface system 905 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 905 may include one or more interfaces between the control system 910 and a memory system, such as the optional memory system 915 shown in Figure 9A. However, the control system 910 may include a memory system in some instances. The interface system 905 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0116] The control system 910 may, for example, include a general purpose single- or multi- chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0117] In some implementations, the control system 910 may reside in more than one device. For example, in some implementations a portion of the control system 910 may reside in a device within one of the environments depicted herein and another portion of the control system 910 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 910 may reside in a device within one environment and another portion of the control system 910 may reside in one or more other devices of the environment. For example, a portion of the control system 910 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 910 may reside in another device that is implementing the cloud- based service, such as another server, a memory device, etc. The interface system 905 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0118] In some implementations, the control system 910 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 910 may be configured for implementing methods of determining a probability an echo is present in an audio signal, performing echo suppression based on the determined probability, or the like.D23133EP
[0119] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 915 shown in Figure 9A and / or in the control system 910. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, filter audio signals to reduce noise, determine filter coefficients to perform such filtering, etc. The software may, for example, be executable by one or more components of a control system such as the control system 910 of Figure 9A.
[0120] In some examples, the apparatus 900 may include the optional microphone system 920 shown in Figure 9A. The optional microphone system 920 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 900 may not include a microphone system 920. However, in some such implementations the apparatus 900 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 910. In some such implementations, a cloud-based implementation of the apparatus 900 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 910.
[0121] According to some implementations, the apparatus 900 may include the optional loudspeaker system 925 shown in Figure 9A. The optional loudspeaker system 925 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 900 may not include a loudspeaker system 925. In some implementations, the apparatus 900 may include headphones. Headphones may be connected or coupled to the apparatus 900 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0122] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor,D23133EP programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0123] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0124] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.
[0125] Figure 9B illustrates a schematic block diagram of an example device architecture 901 (in this example, an apparatus 901) that may be used to implement various aspects of the present disclosure. The apparatus 901 of Figure 9B is an instance of the apparatus 900 of Figure 9A. Architecture 901 includes but is not limited to servers and client devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of Figure 8. As shown, the architecture 901 includes central processing unit (CPU) 941, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 942 or a program loaded from, for example, storage unit 948 to random access memory (RAM) 943. The CPU 941 may be, for example, an electronic processor 941. In these examples, the CPU 941 is an instance of the control system 910 of Figure 9A and the ROM 942D23133EP and RAM 943 are instances of the memory system 915. In RAM 943, the data required when CPU 941 performs the various processes is also stored, as required. CPU 941, ROM 942, and RAM 943 are connected to one another via bus 944. Input / output (I / O) interface 945 is also connected to bus 944. The bus 944 and the I / O) interface 945 are instances of the interface system 905 of Figure 9A.
[0126] The following components are connected to I / O interface 945: input unit 946, that may include a keyboard, a mouse, or the like; output unit 947 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 948 including a hard disk, or another suitable storage device; and communication unit 949 including a network interface card such as a network card (e.g., wired or wireless).
[0127] In some implementations, input unit 946 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0128] In some implementations, output unit 947 include systems with various number of speakers. Output unit 947 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0129] In some embodiments, communication unit 949 is configured to communicate with other devices (e.g., via a network). Drive 950 is also connected to I / O interface 945, as required. Removable medium 951, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 950, so that a computer program read therefrom is installed into storage unit 948, as required. A person skilled in the art would understand that although apparatus 901 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0130] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 949, and / or installed from the removable medium 951, as shown in Figure 9B.D23133EP
[0131] Figure 9C illustrates a schematic block diagram of an example CPU 941 implemented in the device architecture 901 of Figure 9B that may be used to implement various aspects of the present disclosure. The CPU 941 includes an electronic processor 960 and a memory 961. The electronic processor 960 is electrically and / or communicatively connected to the memory 961 for bidirectional communication. The memory 961 stores encoding software 962 and decoding software 963. The memory 961 may be, for example, a ROM, a RAM, or another non-transitory computer readable medium. The electronic processor 960 may implement the encoding software 962 stored in the memory 961 to perform, among other things, the method 800 of Figure 8. Additionally, the electronic processor 960 may implement the decoding software 963 stored in the memory 961 to perform, among other things, the methods that are described with reference to any or all of Figure 8.
[0132] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 941 in combination with other components of Figure 9B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0133] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0134] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium mayD23133EP be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0136] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.
[0137] EEE 1. A method for suppressing echoes in audio signals, the method comprising: receiving a reference signal indicative of an echo signal associated with an audio capture and playback device (802); receiving an audio signal (804) from the audio capture and playback device (100); determining a probability of echo present in the audio signal based at least in part on a comparison of spatial characteristics of the reference signal and spatial characteristics of the audio signal (806); andD23133EP performing echo suppression on the audio signal based at least in part on the probability of echo present in the audio signal (808).
[0138] EEE 2. The method of EEE 1, further comprising: determining a power vector associated with the reference signal, the power vector associated with the reference signal being indicative of a direction of arrival of the reference signal; and determining a power vector associated with the audio signal, the power vector associated with the audio signal being indicative of a direction of arrival of the audio signal, wherein determining the probability of echo comprises comparing the power vector associated with the reference signal and the power vector associated with the audio signal.
[0139] EEE 3. The method of EEE 2, wherein comparing the power vector associated with the reference signal and the power vector associated with the audio signal comprises determining a cosine similarity.
[0140] EEE 4. The method of any one of EEE 1-3, wherein determining the probability of echo comprises applying a beamforming technique to the audio signal.
[0141] EEE 5. The method of EEE 4, wherein applying the beamforming technique comprises: generating a first beamformed signal by suppressing ambient sound in the audio signal while preserving echo signal in the audio signal; generating a second beamformed signal by suppressing the echo signal in the audio signal while preserving the ambient sound in the audio signal; determining a first power associated with the first beamformed signal and a second power associated with the second beamformed signal; and determining the probability of echo based on the first power and the second power.
[0142] EEE 6. The method of any one of EEE 1-5, wherein performing echo suppression on the audio signal comprises modifying an adaptation rate of an acoustic echo cancellation (AEC) system based on the probability of echo present in the audio signal.
[0143] EEE 7. The method of EEE 6, wherein the adaptation rate is proportional to the probability of echo present in the audio signal.
[0144] EEE 8. The method of any one of EEE 1-7, wherein performing echo suppression on the audio signal comprises suppressing residual echo remaining in the audio signal after the audio signal has been processed by an adaptive echo cancellation (AEC) system.D23133EP
[0145] EEE 9. The method of EEE 8, wherein suppressing residual echo comprises determining a suppression gain based at least in part on the probability of echo present in the audio signal.
[0146] EEE 10. The method of any one of EEE 1-9, further comprising augmenting the audio signal after performing echo suppression to compensate for over-suppression of echo in the audio signal.
[0147] EEE 11. The method of EEE 10, wherein augmenting the audio signal comprises selecting an infill signal from one or more of: a microphone channel selected as least affected by echo; an output of a beamforming technique; or signal from frequency bands of the audio signal selected as least affected by echo.
[0148] EEE 12. The method of EEE 11, further comprising scaling the infill signal based at least in part on a probability of residual echo after performing the echo suppression and a level of content in the audio signal.
[0149] EEE 13. The method of EEE 12, wherein the level of content in the audio signal is determined based on a smoothing factor that considers the probability of residual echo and a cross- correlation between the reference signal and the audio signal after performing the echo suppression.
[0150] EEE 14. The method of any one of EEE 1-6, wherein the audio signal comprises n channels of microphone signals, wherein the probability of echo present in the audio signal is a mono echo probability indicating a likelihood of echo across all n channels, and wherein the method further comprises: expanding the mono echo probability to the n channels to obtain an echo probability for each of the n channels of microphone signals; and determining an adaptation rate for each of the n channels based on the echo probability for the respective channel.
[0151] EEE 15. The method according to EEE 14, wherein determining the echo probability for each of the n channels of microphone signals comprises multiplying the mono echo probability with an echo power for the given channel, wherein the echo power is determined from a covariance between the n channels for the reference signal.
[0152] EEE 16. The method according to any one of EEE 14-15, wherein the adaption rate for each of the n channels is determined to be equal to the echo probability for the respective channelD23133EP on a condition that an update flag is set, and otherwise to be equal to zero, wherein the update flag is set responsive to an echo reference power exceeding a threshold.
[0153] EEE 17. An apparatus configured for implementing the method of any one of EEE 1-16.
[0154] EEE 18. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of EEE 1-16.
Claims
D23133EP CLAIMS 1. A method for suppressing echoes in audio signals, the method comprising: receiving a reference signal indicative of an echo signal associated with an audio capture and playback device (802); receiving an audio signal (804) from the audio capture and playback device (100); determining a probability of echo present in the audio signal based at least in part on a comparison of spatial characteristics of the reference signal and spatial characteristics of the audio signal (806); and performing echo suppression on the audio signal based at least in part on the probability of echo present in the audio signal (808), wherein performing echo suppression on the audio signal comprises modifying an adaptation rate of an acoustic echo cancellation, AEC, system based on the probability of echo present in the audio signal.
2. The method of claim 1, further comprising: determining a power vector associated with the reference signal, the power vector associated with the reference signal being indicative of a direction of arrival of the reference signal; and determining a power vector associated with the audio signal, the power vector associated with the audio signal being indicative of a direction of arrival of the audio signal, wherein determining the probability of echo comprises comparing the power vector associated with the reference signal and the power vector associated with the audio signal.
3. The method of claim 2, wherein comparing the power vector associated with the reference signal and the power vector associated with the audio signal comprises determining a cosine similarity.
4. The method of any one of claim 1-3, wherein determining the probability of echo comprises applying a beamforming technique to the audio signal.
5. The method of claim 4, wherein applying the beamforming technique comprises: generating a first beamformed signal by suppressing ambient sound in the audio signal while preserving echo signal in the audio signal; generating a second beamformed signal by suppressing the echo signal in the audio signal while preserving the ambient sound in the audio signal;D23133EP determining a first power associated with the first beamformed signal and a second power associated with the second beamformed signal; and determining the probability of echo based on the first power and the second power.
6. The method of any one of claims 1-5, wherein the adaptation rate is proportional to the probability of echo present in the audio signal.
7. The method of any one of claims 1-6, wherein the audio signal comprises n channels of microphone signals, wherein the probability of echo present in the audio signal is a mono echo probability indicating a likelihood of echo across all n channels, and wherein the method further comprises: expanding the mono echo probability to the n channels to obtain an echo probability for each of the n channels of microphone signals; and determining an adaptation rate for each of the n channels based on the echo probability for the respective channel.
8. The method according to claim 7, wherein determining the echo probability for each of the n channels of microphone signals comprises multiplying the mono echo probability with an echo power for the given channel, wherein the echo power is determined from a covariance between the n channels for the reference signal.
9. The method according to any one of claims 7-8, wherein the adaption rate for each of the n channels is determined to be equal to the echo probability for the respective channel on a condition that an update flag is set, and otherwise to be equal to zero, wherein the update flag is set responsive to an echo reference power exceeding a threshold.
10. The method of any one of claims 1-9, wherein performing echo suppression on the audio signal further comprises suppressing residual echo remaining in the audio signal after the audio signal has been processed by the adaptive echo cancellation (AEC) system.
11. The method of claim 10, wherein suppressing residual echo comprises determining a suppression gain based at least in part on the probability of echo present in the audio signal.
12. The method of any one of claims 1-11, further comprising augmenting the audio signal after performing echo suppression to compensate for over-suppression of echo in the audio signal,D23133EP wherein, optionally, augmenting the audio signal comprises selecting an infill signal from one or more of: a microphone channel selected as least affected by echo; an output of a beamforming technique; or signal from frequency bands of the audio signal selected as least affected by echo.
13. The method of claim 12, further comprising scaling the infill signal based at least in part on a probability of residual echo after performing the echo suppression and a level of content in the audio signal, wherein, optionally, the level of content in the audio signal is determined based on a smoothing factor that considers the probability of residual echo and a cross-correlation between the reference signal and the audio signal after performing the echo suppression.
14. An apparatus configured for implementing the method of any one of claims 1-13.
15. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of claims 1-13.
Citation Information
Patent Citations
Echo suppression method and device, electronic equipment and storage medium
CN115440236A
Multi-channel echo cancellation and noise suppression
EP2992668B1
Downlink activity and double talk probability detector and method for an echo canceler circuit
US20050129226A1
Controlling Operational Characteristics of Acoustic Echo Canceller
US20160127527A1