Audio processing system, semiconductor device and method for acoustic echo cancellation
By using spatial filtering logic and linear adaptive filters in audio processing systems, the problem of difficulty in eliminating nonlinear acoustic echoes in the prior art is solved, and more efficient echo cancellation and system performance improvement is achieved.
Patent Information
- Application Number
- CN202080067039.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-27
- Filing Date
- 2020-09-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-08
AI Technical Summary
In existing audio processing systems, linear filters are difficult to effectively eliminate nonlinear acoustic echoes, resulting in residual nonlinear echoes in the target speech signal, limiting system performance.
Spatial filtering logic is used to receive reference signals and multi-channel microphone signals, generate spatially filtered signals, and use linear adaptive filters to eliminate linear and nonlinear echoes carried in the signal.
The spatial filtering technology effectively eliminates nonlinear echoes, reduces the computational complexity, improves the accuracy of the output signal, and tracks the robustness of nonlinear changes over time in the system.
Smart Images

Figure CN114450973B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is an international application of U.S. Patent Application No. 16 / 585,750, filed on September 27, 2019, the entire content of which is incorporated herein by reference in its entirety. Technical field
[0003] This disclosure relates to signal processing in an audio processing system. Background art
[0004] In audio processing systems such as smart speakers, hands - free telephones, and speech recognition systems, the use of high - power speakers is growing rapidly. In such audio processing systems, acoustic coupling typically occurs between the speaker and the microphone during playback and / or voice interaction. For example, the audio signal played by the speaker is captured by the microphone in the system. Audio signals typically generate acoustic echoes when they propagate in a confined space (e.g., inside a room, a vehicle, etc.), but such acoustic echoes are undesirable because they may dominate the target speech signal.
[0005] To eliminate unwanted acoustic echoes, audio processing systems typically use an acoustic echo canceller (AEC) with a linear filter to estimate the room impulse response (RIR) transfer function, which characterizes the propagation of acoustic signals in a confined space. However, the estimation model used by the linear filter in such an AEC is not suitable for modeling any non - linearities in the captured acoustic signals because such non - linearities have non - uniform origins, may change over time, and are computationally very expensive and difficult to estimate. Failure to properly cancel such non - linearities can result in residual non - linear echoes in the target speech signal, which may severely limit the performance of any system (e.g., such as a speech recognition system) that processes the target signal. Summary of the invention
[0006] According to one aspect of the present disclosure, an audio processing system is disclosed. The audio processing system includes: a loudspeaker configured to receive a reference signal; a microphone array configured to provide a multi-channel microphone signal, the multi-channel microphone signal including both linear echo and non-linear echo; spatial filtering logic configured to receive the reference signal and the multi-channel microphone signal and generate a spatially filtered signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; acoustic echo canceller (AEC) logic configured to at least: receive the spatially filtered signal; and apply a linear adaptive filter using the spatially filtered signal to generate a cancellation signal, the cancellation signal estimating both the linear echo and the non-linear echo of the multi-channel microphone signal; and a logic block configured to receive the cancellation signal and generate an output signal based at least on the cancellation signal.
[0007] According to one aspect of the present disclosure, a semiconductor device for audio processing is disclosed. The semiconductor device includes a digital signal processor (DSP) configured to: receive a reference signal sent to a loudspeaker; receive a multi-channel microphone signal from a microphone array, wherein the multi-channel microphone signal includes both linear echo and non-linear echo; apply a spatial filter using the reference signal and the multi-channel microphone signal to generate a spatially filtered signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; apply a linear adaptive filter using the spatially filtered signal to generate a cancellation signal, the cancellation signal estimating both the linear echo and the non-linear echo of the multi-channel microphone signal; and generate an output signal based at least on the cancellation signal.
[0008] According to one aspect of the present disclosure, a method for acoustic echo cancellation is disclosed. The method includes: receiving a reference signal sent to a loudspeaker; receiving a multi-channel microphone signal from a microphone array acoustically adjacent to the loudspeaker, wherein the multi-channel microphone signal includes both linear echo and non-linear echo; generating, by a processing device, a spatially filtered signal by applying a spatial filter using the reference signal and the multi-channel microphone signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; generating, by the processing device, a cancellation signal by applying a linear adaptive filter using the spatially filtered signal, wherein the cancellation signal estimates both the linear echo and the non-linear echo of the multi-channel microphone signal; and generating, by the processing device, an output signal based at least on the cancellation signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figures 1A to 1CIllustrates an example system for non - linear acoustic echo cancellation according to some embodiments.
[0010] Figures 2A to 2C Illustrates a flowchart of an example method for non - linear acoustic echo cancellation according to some embodiments.
[0011] Figures 3A to 3B Illustrates a graph reflecting a simulation study of the described techniques for non - linear acoustic echo cancellation.
[0012] Figure 4 Illustrates a schematic diagram of an example audio processing device according to some embodiments.
[0013] Figure 5 Illustrates a schematic diagram of an example host device according to some embodiments. Detailed Description
[0014] The following description sets forth examples of numerous specific details, such as specific systems, components, methods, etc., in order to provide a good understanding of various embodiments of the described techniques for non - linear acoustic echo cancellation. However, it will be apparent to those skilled in the art that at least some embodiments may be practiced without these specific details. In other instances, well - known components, elements, or methods are not described in detail but are presented in a simple block diagram format in order to avoid unnecessarily obscuring the subject matter described herein. Thus, the specific details set forth below are merely exemplary. Specific implementations may differ from these exemplary details and still be considered within the spirit and scope of the invention.
[0015] References in the specification to "embodiments", "one embodiment", "example embodiments", "some embodiments", and "various embodiments" mean that a particular feature, structure, step, operation, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Moreover, the phrases "embodiments", "one embodiment", "example embodiments", "some embodiments", and "various embodiments" appearing throughout the specification do not necessarily all refer to the same embodiment. References to "cancel", "cancelling", and other verb derivatives thereof mean to completely or at least substantially remove an unwanted signal (e.g., such as a linear or non - linear echo) from another signal (e.g., such as an output signal).
[0016] This description includes references to the accompanying drawings, which form a part of the detailed description. The drawings illustrate diagrams in accordance with exemplary embodiments. These embodiments, which may also be referred to herein as "examples," are described in sufficient detail to enable those skilled in the art to practice the embodiments of the claimed subject matter described herein. Embodiments may be combined, other embodiments may be utilized, or structural, logical, and electrical changes may be made without departing from the scope and spirit of the claimed subject matter. It should be understood that the embodiments described herein are not intended to limit the scope of the subject matter, but rather to enable those skilled in the art to practice, make, and / or use the subject matter.
[0017] Various embodiments of techniques for performing non-linear echo cancellation in a device that provides audio processing are described herein. Examples of such devices include, but are not limited to, personal computers (e.g., laptop computers, notebook computers, etc.), mobile computing devices (e.g., tablets, tablet computers, etc.), conference telephone devices (e.g., speakerphones, etc.), mobile communication devices (e.g., smart phones, etc.), smart speakers, printed circuit board (PCB) modules configured for audio processing, system-on-a-chip (SoC) semiconductor devices and multi-chip semiconductor packages, Internet of Things (IoT) wireless devices, and other similar electronic devices, computing devices, and on-chip devices for audio processing.
[0018] Generally, an echo is a signal generated by transforming an acoustic and / or audio signal through the transfer function of components in an audio system. Such an echo is typically unwanted because it may dominate the target speech signal. To remove the unwanted echo, a front-end audio processing system typically uses an acoustic echo canceller (AEC) to remove the echo signal from the target audio signal before the target audio signal is sent to a backend system. The backend system, which may run on the cloud or on a local computer, requires the audio signals it receives to be as clean as possible. For example, a microphone coupled to the front-end system receives acoustic (sound) waves and converts them into an analog audio signal, which is then digitized. However, the received sound waves may have been interfered with by nearby devices (e.g., an on television, etc.) or acoustic echoes from speakers. For example, a person (whose speech needs to be recognized) may be speaking while a speaker is playing music or other multimedia content, and such playback is also captured by the microphone as an echo along with the speaker's speech.
[0019] Since the transfer functions of the components of an audio system can be linear and / or non-linear, an audio system typically generates both linear echoes and non-linear echoes. A linear echo is a signal generated by transforming an acoustic / audio signal through a linear transfer function, the output of which is a linear combination of its input signal. On the other hand, a non-linear (NL) echo is a signal generated by transforming an acoustic / audio signal through a non-linear transfer function, the output of which is not a linear combination of its input signal. A non-linear transfer function does not satisfy one or more of the conditions of linearity, which require that the output level be proportional to the input level (homogeneity), and that the response caused by two or more input signals be the sum of the responses that would have been caused by each input signal individually (additivity). Thus, the echoes in a typical audio system are signals generated by transforming an acoustic / audio signal through the linear and non-linear transfer functions of the components in the system, the linear and non-linear transfer functions including, for example, the transfer functions of the loudspeakers, power amplifiers, and microphones of the system, as well as the RIR transfer function that characterizes the propagation of acoustic signals in a confined space and / or the physical environment of the system.
[0020] The AEC in a typical audio system can only access the linear reference signal, so it uses a linear filter to remove the linear echo from the target audio signal. However, the estimation model used by the linear filter in such an AEC is not suitable for modeling any non-linearity in the captured acoustic signal, because such non-linearity has a non-uniform origin, may change over time, and is computationally very expensive and difficult to estimate. Thus, in a typical audio system with a linear AEC, any non-linear echo generated by the system remains in the target signal.
[0021] For example, a typical audio processing system can have various non - linearities with different non - linear transfer functions, and thus the combined non - linear echo in such a system can have multiple origins. The transfer functions of active components (e.g., transistors, amplifiers, power supplies, etc.) and passive components (e.g., speaker components such as cones and diaphragms, etc.) in the system may have non - linearities, which may be sources of signal distortion. When picked up by a microphone in the system, such non - linear signal distortion can be the cause of unwanted non - linear echo. Compared with linear distortion (which is the expected distortion caused by a linear transfer function), non - linear distortion is unexpected because they are (at least partially) caused by the current physical condition of the speaker - for example, fatigue of speaker components, wear of speaker assemblies, and the condition of speaker cones. The physical condition of the speaker will necessarily deteriorate over time, which further changes the non - linear distortion generated by the components of the speaker. In addition, operating the speaker at or beyond its sound limit may also cause non - linear distortion due to unpredictable vibrations of the speaker assembly and / or its components.
[0022] In addition, AEC processing is computationally very expensive. This problem is more severe in systems with limited computational capabilities, such as embedded systems (e.g., SoC), IoT devices that provide front - end processing for back - end automatic speech recognition systems (e.g., such as Amazon Alexa, Google Home, etc.), and edge devices that do not have the ability to perform a large amount of computation (e.g., the entry point into an IoT cloud - based service). Moreover, in some operating environments, echo cancellation is implemented in smart speaker systems that need to identify and respond to voice commands in real - time. Therefore, in such operating environments, the computational cost of AEC processing is an important factor in the system response time.
[0023] The acoustic echo generated when the signal broadcast by the speaker is captured by the microphone can be represented according to Equation (1) below:
[0024] d(n) = h(n) T x(n) (1)
[0025] where h(n) is the vector of the impulse response between the speaker and the microphone, (·) T is the transpose operator, x(n) is the vector of the reference (e.g., speaker) signal, d(n) is the vector of the acoustic echo signal captured by the microphone, and n is the time index. In a conventional audio processing system, the purpose of linear AEC is to obtain an acoustic echo estimation signal using a linear filter with coefficients w(n) according to Equation (2) below
[0026]
[0027] So that according to Equation (3), the mean square error is minimized over time:
[0028]
[0029] where e(n) is the output signal after echo cancellation, y(n) is the microphone signal (which can include other signals such as speech captured along with the acoustic echo), and E[·] is the expectation (average) operator.
[0030] The overall challenge of non-linear AEC is that the acoustic coupling between the speaker and the microphone cannot be simply modeled linearly by Equation 2. A conventional way to address this challenge is to generalize the problem by modeling the non-linearity using some closed-form function f NL (e.g., Volterra filter, Hammerstein filter, neural network, etc.) according to Equation (4) below, for example:
[0031]
[0032] However, due to high complexity (e.g., O(N 2 ) compared to O(N) of the linear adaptive filter, which may be prohibitive for real-time systems), the manifestation of local minima (e.g., it does not always eliminate echoes), slow convergence (e.g., it has a slow echo cancellation rate), low accuracy (e.g., the quality of its output signal after echo cancellation is poor), and numerical instability (e.g., it has limited numerical precision, which may cause the system to drift over time and become unstable), this conventional method often becomes computationally impractical. This conventional method also suffers from a lack of knowledge about the non-linearity in the actual audio system and the variation of the non-linearity over time (e.g., the physical condition of the speaker), which makes it almost impossible to design a practical solution using a model that takes into account all possible non-linearities during the lifetime of the system.
[0033] Another conventional method is to decouple the linear and non-linear echo cancellation processes. For example, this conventional method involves applying a preprocessing filter f PRE to the reference signal before applying the reference signal to the linear filter of the AEC to best match the effect of a specific expected non-linearity, e.g., according to Equation (5) below:
[0034]
[0035] where is the transformed reference signal used by the linear filter to obtain the echo estimate according to equation (2) above. While this conventional method may be computationally feasible for some applications, it is not robust (e.g., it cannot account for non-linear variations in the actual system), and cannot correctly model non-linear time variability (e.g., which is caused by the gradual degradation of the physical condition of the speaker).
[0036] To address these and other drawbacks of the conventional method, the techniques described herein provide for removing non-linear echo by applying a spatial filter to the speaker reference signal. Spatial filtering can be specifically performed in the direction of the speaker and / or in the direction of the talker, provided that the target talker and the speaker are sufficiently spatially separated (e.g., this is the case in most situations if not all applications using a microphone array). The spatially filtered signal (e.g., having reduced amplitude and frequency in some of its parts) is then used as the reference signal provided to the AEC with the linear filter. In this way, the techniques described herein provide low computational complexity (e.g., since a linear AEC is used to cancel non-linear echo), while providing robustness (e.g., since non-linearities in the system can be tracked over time) and improved accuracy of the output signal (e.g., compared to the conventional method). Additionally, when the non-linearities in the system change over time, a system according to the described techniques does not need to be retuned as required under the conventional method.
[0037] Generally, spatial filtering is a transformation of the input multi-channel signal measured at different spatial positions such that its output signal depends only on its input signal. Examples of spatial filters include, but are not limited to, re-reference filters, surface Laplacian filters, independent component analysis (ICA) filters, and common spatial pattern (CSP) filters. In practice, spatial filtering is typically incorporated into an audio processing system using a beamformer. A beamformer (BF) is a signal processing mechanism that steers the spatial responses of multiple microphones (e.g., in a microphone array) towards a target audio source, and thus naturally measures or otherwise collects all the parameters required to construct the spatial filter.
[0038] According to the techniques for non-linear echo cancellation described herein, the spatial filter f SF is used to capture the effects of non-linearities in the audio processing system. The spatial filter f SF generates the spatially filtered reference signal, for example, according to equation (6) below
[0039]
[0040] where is a spatially filtered signal that is directed towards a loudspeaker and used for adaptive filtering (which includes both a linear reference signal and a non-linear reference signal), is a multi-channel signal received from multiple microphones, and n is a time index.
[0041] In some embodiments, the techniques described herein may be implemented in an audio processing system having a single AEC applied to one of the multiple microphone signals. In these embodiments, the spatially filtered signal is used as a reference signal provided to an AEC having a linear filter to obtain, for example, an echo estimation signal according to the following equation (7)
[0042]
[0043] such that the mean square error is minimized over time, for example, according to the following equation (8):
[0044]
[0045] where i = 1, 2, …, N and N is the number of microphones, y i (n) is the microphone signal selected from the i-th microphone for processing, w(n) is the vector of linear filter coefficients of the AEC, is the spatially filtered reference signal (e.g., as generated according to the above equation (6)), n is the time index, and is the target speech signal from the i-th microphone provided as the echo-canceled output signal.
[0046] In some embodiments, the techniques described herein may be implemented in an audio processing system having multiple linear AECs or a multi-instance linear AEC with one linear AEC instance provided for each microphone. In these embodiments, the same spatial filter f SF is used to direct towards the loudspeaker to extract the non-linear reference signal and towards the main speaker to extract the target speech estimation signal The speech estimation signal can be obtained by applying the same spatial filter f SF (but with different coefficients) to the multi-channel output of the AEC, for example, according to the following equation (9)
[0047]
[0048] since the mean square error is minimized over time, for example, according to the following equation (10):
[0049]
[0050] where \(i = 1, 2, \ldots, N\) and \(N\) is the number of microphones, is the multi-channel output signal from a multi-instance (or multiple) linear AEC, is the echo estimate vector, and \(y\) i (n) is the microphone signal from the \(i\)-th microphone, and \(w\) i (n) is the vector of linear filter coefficients for the linear filter associated with the \(i\)-th microphone, is the spatially filtered reference signal (e.g., generated as per equation (6) above), \(n\) is the time index, and (e.g., generated as per equation (9) above) is the target speech estimate signal provided as the echo-canceled output signal.
[0051] In some embodiments, the techniques described herein can be implemented in an audio processing system having a single linear AEC applied to the output from a spatial filter \(f\) SF . In these embodiments, a spatial filter \(f\) with appropriate filter coefficients is used to extract the spatially filtered reference signal SF (e.g., generated as per equation (6) above) and the spatially filtered microphone signal . The spatially filtered reference signal is directed towards the speaker to extract any non-linear reference signal. The spatially filtered microphone signal includes both the spatially amplified speech estimate signal and the attenuated echo estimate signal (e.g., ), and can be generated according to equation (11) below by using appropriate spatial filter coefficients: such that the mean square error is minimized over time, e.g., according to equation (12) below:
[0052]
[0053] where
[0054]
[0055] where is the spatially filtered microphone signal, and \(w\) i (n) is the vector of linear filter coefficients of the linear AEC, is the spatially filtered reference signal (e.g., generated as per equation (6) above), \(n\) is the time index, and is a target speech estimation signal provided as an output signal with echo cancellation.
[0056] In some embodiments, the techniques described herein can be used in systems with an adaptive beamformer that will naturally be able to capture and track non-linear variations in the system. In these embodiments, the signal output from the BF block includes both the linear and non-linear components of the reference signal, and thus the linear AEC can cancel the non-linear part of the echo. Additionally, the techniques described herein are not limited to adaptive beamforming, but can be used with other spatial filtering techniques - for example, switched beamforming, (blind or semi-blind) source separation, etc.
[0057] Figures 1A to 1C Systems 100A to 100C for non-linear acoustic echo cancellation according to example embodiments are shown respectively. In some embodiments (e.g., such as a conference phone device), the components of each of systems 100A to 100C can be integrated into the same housing as a stand-alone device. In other embodiments (e.g., a smart speaker system), the components of each of systems 100A to 100C can be separate elements coupled by one or more networks and / or communication lines. In other embodiments, the components of each of systems 100A to 100C can be arranged in a spatially separated fixed speaker-microphone geometry that provides speakers to potential speakers. Thus, Figures 1A to 1C the systems 100A to 100C in are considered in an illustrative sense rather than a limiting sense.
[0058] In Figures 1A to 1C similar reference numerals refer to similar components. Thus, Figures 1A to 1C each of the systems 100A to 100C in includes a speaker-microphone assembly 110 coupled to an audio processing device 120, which is coupled to a host 140. The audio processing device 120 includes spatial filtering logic 124, AEC logic 126, and adder logic 128. As used herein, "logic" refers to a hardware block having one or more circuits, the one or more circuits including various electronic components configured to process analog and / or digital signals and perform one or more operations in response to control signals and / or firmware instructions executed by a processor or its equivalent. Examples of such electronic components include, but are not limited to, transistors, diodes, logic gates, state machines, microcoded engines, and / or other circuit blocks and analog / digital circuitry that can be configured to control the hardware in response to control signals and / or firmware instructions.
[0059] The speaker - microphone assembly 110 includes one or more speakers 112 and a microphone array 114, which are arranged in acoustic proximity such that the microphone array can detect sound waves from a desired sound source (e.g., human speech) and sound waves from an undesired sound source (e.g., an acoustic echo 113 such as from the speaker 112). As used herein, a "speaker" refers to an electro - acoustic speaker device configured to transform an electrical signal into an acoustic / sound wave. The speaker 112 is configured to receive an analog audio signal from an audio processing device 120 and emit the audio signal as a sound wave. The microphone array 114 includes a plurality of microphones configured to receive sound waves from various sound sources and transform the received sound waves into an analog audio signal that is sent to the audio processing device 120. In some embodiments (e.g., smart phones), the speaker 112 and the microphone array 114 can be integrally formed as the same speaker - microphone assembly 110. In some embodiments (e.g., conference phone devices), the speaker 112 and the microphone array 114 can be separate components arranged on a common substrate (e.g., a PCB) that is mounted inside or on the housing of the speaker - microphone assembly 110. In yet another embodiment, the speaker - microphone assembly 110 may not have a housing but can be formed by the acoustic proximity of the speaker 112 and the microphone array 114.
[0060] The audio processing device 120 includes spatial filtering logic 124, AEC logic 126, and adder logic 128. In some embodiments, the audio processing device 120 can be a single - chip integrated circuit (IC) device fabricated on a semiconductor die or a single - chip IC fabricated as a SoC. In other embodiments, the audio processing device 120 can be a multi - chip module encapsulated in a single semiconductor package or arranged or mounted on a common substrate such as a PCB with multiple semiconductor packages. In some embodiments, the spatial filtering logic 124, AEC logic 126, and adder logic 128 can be implemented as hardware circuitry within a digital signal processor (DSP) of the audio processing device 120. In various embodiments, the audio processing device 120 can include additional components (not shown), such as audio input / output (I / O) logic, a central processing unit (CPU), memory, and one or more interfaces for connecting to a host 140.
[0061] In some embodiments, the spatial filtering logic 124 may implement or may be implemented as part of the BF logic that steers the spatial response of the microphone array 114 toward a target audio source. For example, such BF logic may apply time delay compensation to the digital signals from each microphone in the microphone array 114 to compensate for the relative time delays between the microphone signals that may result from the position of the sound source relative to each microphone. The BF logic may also be configured to attenuate the digital signals from some of the microphones, amplify the digital signals from other microphones, and / or alter the directivity of the digital signals from some or all of the microphones. In some embodiments, such BF logic may also use the signals received from the sensors in the microphone array 114 to track a moving speaker and adjust the digital signals from each microphone accordingly. In this way, the BF logic measures or otherwise collects the parameters required to operate one or more instances of the spatial filtering logic configured to apply one or more spatial filters (or instances thereof) to its input signals.
[0062] In accordance with the techniques described herein, the spatial filtering logic 124 is configured to apply spatial filters (e.g., Figure 1A 124a in Figure 1B and Figure 1C 124a-1 and 124a-2 in Figures 1A to 1C ) to the multi-channel microphone signals received from the microphone array 114. The spatial filtering logic 124 is configured to generate a spatially filtered signal targeted at a particular direction toward a particular audio source. For example, in an embodiment of Figures 1A to 1C , the spatial filtering logic 124 may perform spatial filtering in the direction of the speaker 112 to generate a spatially filtered signal that includes both linear echo and non-linear echo. In other embodiments (e.g., Figure 1C ), the spatial filtering logic 124 may additionally perform spatial filtering in the direction of the speaker (e.g., based on the multi-channel output signals) to generate a spatially filtered signal that includes a speech estimation signal.
[0063] According to the techniques described herein, the AEC logic 126 includes linear filter logic 126a for generating an echo estimation signal that is removed from the output signal that is ultimately sent to the host 140. In some embodiments, the linear filter logic 126a implements the following linear adaptive filter, whose output is a linear combination of its inputs and whose transfer function is controlled by variable parameters that can be adjusted during operation based on the output signal 129 generated by the adder logic 128. (However, it should be noted that various embodiments of the techniques described herein may use various other types of linear filters). Generally, adaptive filtering is a technique of continuously adjusting the filter coefficients of the AEC to reflect a changing acoustic environment (e.g., when different speakers start speaking, when the microphone or speaker is physically moved, etc.) so as to achieve the best possible filtered output (e.g., by minimizing the residual echo energy over time according to equation (3) above). Adaptive filtering can be implemented in a sampled manner in the time domain or in a block-wise manner in the frequency domain over time. A typical implementation of a linear adaptive filter (e.g., such as the linear filter logic 126a) may use background and foreground filtering. Background-foreground filtering is an adaptive filtering technique that involves two separate adaptive filters ("background" and "foreground") that are combined to maximize system performance. The background filter is designed to actively and quickly adapt and eliminate as much echo as possible in a short period of time at the expense of reduced noise stability, while the foreground filter adjusts cautiously from a long-term perspective at the expense of a slow convergence rate to provide a stable and optimal output. In this way, the foreground filter can maintain convergence even in the presence of noise, while the background filter can capture any rapid changes and dynamics in the acoustic environment. In practice, a linear adaptive filter with background-foreground filtering is typically required to handle barge-in scenarios and double-talk scenarios in a robust manner. "Double-talk" is a scenario that occurs during a conference call when the local / proximal speaker and the remote / distal speaker are speaking simultaneously, causing the local voice signal and the remote voice signal to be captured simultaneously by the local microphone. "Barge-in" is a scenario similar to double-talk, except that the real-time remote speaker is replaced by a device / machine that may be playing back the captured voice signal itself or a multimedia signal such as music.
[0064] The adder logic 128 performs a digital summation on its input digital signals and generates an output signal 129 (e.g., in Figure 1A system 100A of Figure 1C system 100C) or a multi-channel output signal 129a (in Figure 1BLogic blocks in system 100B). Digital summation involves adding and / or subtracting two or more signals using element-by-element indexing - for example, adding the nth sample of one signal to or subtracting it from the nth sample of another signal, and the result represents the nth sample of the output signal.
[0065] Host 140 is coupled to communicate with audio processing device 120. In some embodiments, host 140 may be implemented as a standalone device or as a computing system. For example, host 140 may be implemented on-chip with audio processing device 120 as a SoC device or an IoT edge device. In another example, host 140 may be implemented as a desktop computer, a laptop computer, a conference phone device (e.g., a speakerphone), etc. In other embodiments, host 140 may be implemented in a networked environment as a server computer or server blade communicatively connected to audio processing device 120 via one or more networks.
[0066] In operation, audio processing device 120 receives audio data (e.g., a series of bytes) from host 140. The audio data may represent multimedia playback and / or remote speech. Audio processing device 120 (e.g., one or more of its circuits) ultimately converts the received audio data into a reference signal x(n)111 that is sent to speaker 112. The microphones in microphone array 114 pick up sound waves from proximal speech as well as acoustic echo 113 from speaker 112. The microphones in microphone array 114 convert the received sound waves into corresponding analog audio signals that are sent to audio processing device 120. Audio processing device 120 (e.g., one or more of its circuits) receives the analog audio signals and converts them into a multi-channel digital microphone signal 115, and this multi-channel digital microphone signal 115 is sent to spatial filtering logic 124 for processing according to the techniques described herein. The parameters required for the spatial filter (e.g., such as direction, auto-statistics / cross-channel statistics, optimization function, etc.) may be determined by spatial filtering logic 124, which performs beamforming on the multi-channel microphone signals received from microphone array 114.
[0067] Figure 1A An example system 100A with a single AEC logic 126 is shown. In system 100A, spatial filtering logic 124 applies a spatial filter f SF 124a to the multi-channel microphone signal 115 and generates a spatially filtered signal 125 (e.g., according to equation (6) above). The spatially filtered signal 125 is provided to the AEC logic 126 and carries both linear echo and non-linear echo included in the multi-channel signal 115 - for example, for each time index n, the value sampled from the signal 125 reflects both the linear echo and the non-linear echo picked up by the microphones in the microphone array 114. The AEC logic 126 adaptively calculates the coefficients w(n) of the linear adaptive filter logic 126a. Then the linear adaptive filter logic 126a is applied to the spatially filtered signal 125 (e.g., according to the above equations (7) and (8)) to generate a cancellation signal 127a. The cancellation signal 127a estimates both the linear echo signal and the non-linear echo signal included in the i-th microphone signal y i (n)115a. The cancellation signal 127a and one of the microphone signals of the multi-channel signal 115 (e.g., the i-th) microphone signal is provided as an input to the adder logic 128. The i-th microphone signal y i (n) can be predetermined (e.g., based on the known / fixed arrangement of the speaker 112 relative to the microphone array 114), or can be randomly selected from the channels of the multi-channel microphone signal 115 during operation. The adder logic 128 performs digital summation based on the cancellation signal 127a and based on the selected multi-channel microphone signal y i (n)115 and generates an output signal e(n)129 (e.g., according to the above equation (8)). In fact, the output signal e(n)129 approximates the target speech signal s(n) captured by the i-th microphone (e.g., ). In this way, both the linear echo signal and the non-linear echo signal are eliminated from the output signal e(n)129. Then the output signal e(n)129 is provided to the host 140. Additionally, the output signal e(n)129 is also provided as feedback to the AEC logic 126, and the AEC logic 126 uses this output signal e(n)129 to adaptively calculate the coefficients w(n) of the linear adaptive filter logic 126a.
[0068] In Figure 1A 's implementation, the reference signal x(n)111 is provided to both the speaker 112 and the AEC logic 126. The AEC logic 126 is configured to utilize the reference signal x(n)111 and the spatially filtered signal Both 125. For example, the AEC logic 126 can be configured to use the reference signal x(n) 111 for double-talk detection (DTD). The AEC logic 126 can also be configured to use the spatially filtered signal 125 for its background filter and the reference signal x(n) 111 for its foreground filter, where one of the outputs from the background filter and the output from the foreground filter (e.g., the "best") output is selected to minimize the cancellation of the near-end speech during a double-talk situation.
[0069] Figure 1B An example system 100B including multiple instances of the AEC logic 126 is shown, where one AEC instance is applied to each microphone signal / channel. In system 100B, the spatial filtering logic 124-1 applies an instance 124a-1 of the spatial filter f SF to the multi-channel microphone signal 115 and generates the spatially filtered signal 125 (e.g., according to equation (6) above). The spatially filtered signal 125 is provided to each of the multiple instances of the AEC logic 126 and carries both the linear echo and the non-linear echo included in the multi-channel signal 115. Each instance of the AEC logic 126 adaptively calculates the coefficients w i (n) of its linear adaptive filter 126a, and the linear adaptive filter logic 126a is separately applied to the spatially filtered signal 125 (e.g., according to equation (10) above) to generate the cancellation signal 127b. Thus, the cancellation signal 127b is a multi-channel echo estimation signal that estimates both the linear echo signal and the non-linear echo signal included in the multi-channel microphone signal 115. The multi-channel cancellation signal 127b and the multi-channel microphone signal 115 are provided as inputs to the adder logic 128. The adder logic 128 performs digital summation based on the multi-channel cancellation signal 127b and the multi-channel microphone signal 115 and generates the multi-channel output signal 129a. The spatial filtering logic 124-2 applies an instance 124a-2 of the same spatial filter f SF (e.g., but possibly with different coefficients) to the multi-channel output signal 129a and generates the spatially filtered signal 129 (e.g., according to equation (9) above). In various embodiments, the spatial filtering logic 124-2 may also be configured to receive the reference signal x(n) 111, the multi-channel microphone signal 115 and / or the multi-channel cancellation signal 127b, and use any and / or all of these signals in generating the multi-channel output signal 129a. In fact, the output signal e(n) 129 approximates the target speech signal s(n) captured by the microphones in the microphone array 114 (e.g., ). In this way, both the linear echo signal and the non-linear echo signal are cancelled from the output signal e(n) 129. Then the output signal e(n) 129 is provided to the host 140. Additionally, the multi-channel output signal 129a is also provided as feedback to multiple instances of the AEC logic 126, and the multiple instances use this multi-channel output signal 129a to adaptively calculate the coefficients w i (n) of their respective linear adaptive filter logic 126a.
[0070] In Figure 1B embodiments, the reference signal x(n) 111 is provided to both the speaker 112 and one or more instances of the AEC logic 126. One or more instances of the AEC logic 126 are configured to utilize both the reference signal x(n) 111 and the spatially filtered signal 125. For example, one or more instances of the AEC logic 126 may be configured to use the reference signal x(n) 111 for DTD. Each instance of the AEC logic 126 may also be configured to use the spatially filtered signal 125 for its background filter and the reference signal x(n) 111 for its foreground filter, where one (e.g., "best") output from the output of the background filter and the output of the foreground filter is selected to minimize the cancellation of the proximal speech during a two-way call situation.
[0071] Figure 1C FIG. shows an example system 100C including a single AEC logic 126 applied to the spatial filter output. In system 100C, the spatial filtering logic 124 applies an instance 124a-1 of the spatial filter f SF to the multi-channel microphone signal 115, and generates the spatially filtered signal 125 (e.g., according to equation (6) above). The spatially filtered signal 125 is generated using filter coefficients for the loudspeaker 112 and thus carries both the linear echo and the non-linear echo included in the multi-channel signal 115. Additionally, the spatial filtering logic 124 applies another instance 124a-2 of the same spatial filter f SF to the multi-channel microphone signal 115 and generates a spatially filtered microphone signal 125a (e.g., according to equation (11) above). The spatially filtered microphone signal 125a is generated using filter coefficients for the microphones in the microphone array 114 and thus carries both the spatially amplified speech estimation signal and the attenuated echo estimation signal included in the multi-channel signal 115 (e.g., ). The spatially filtered signal 125 is provided as an input to the AEC logic 126, and the spatially filtered microphone signal 125a is provided as an input to the adder logic 128. The AEC logic 126 adaptively calculates the coefficients w(n) of the linear adaptive filter logic 126a, which is applied to the spatially filtered signal 125 to generate a cancellation signal 127c (e.g., according to equation (12) above). The cancellation signal 127c estimates both the linear echo signal and the non-linear echo signal included in the spatially filtered microphone signal 115. The cancellation signal 127c is provided as an input to the adder logic 128. The adder logic 128 performs a digital summation based on the cancellation signal 127c and based on the spatially filtered microphone signal 125a and generates an output signal e(n) 129 (e.g., according to equation (12) above). In fact, the output signal e(n) 129 approximates the target speech signal s(n) captured by the microphones in the microphone array 114 (e.g., )。In this way, both the linear echo signal and the non-linear echo signal are eliminated from the output signal e(n) 129, and the target speech signal is avoided from being eliminated from the output signal e(n) (e.g., in the case of a two-way call). Then the output signal e(n) 129 is provided to the host 140. Additionally, the output signal e(n) 129 is also provided as feedback to the AEC logic 126, which adaptively calculates the coefficients w(n) of the linear adaptive filter logic 126a using the output signal e(n) 129.
[0072] In Figure 1C the implementation, the reference signal x(n) 111 is provided to both the speaker 112 and the AEC logic 126. The AEC logic 126 is configured to utilize both the reference signal x(n) 111 and the spatially filtered signal 125. For example, the AEC logic 126 can be configured to use the reference signal x(n) 111 for DTD. The AEC logic 126 can also be configured to use the spatially filtered signal 125 for its background filter and the reference signal x(n) 111 for its foreground filter, where one of the outputs from the background filter and the output from the foreground filter (e.g., the "best") output is selected to minimize the cancellation of the near-end speech during a two-way call situation.
[0073] Figures 2A to 2C shows a flowchart of an example method for non-linear acoustic echo cancellation according to the techniques described herein. The operations of the method in Figures 2A to 2C will be described below as being performed by the spatial filtering logic, the AEC logic, and the adder logic (e.g., Figures 1A to 1C the spatial filtering logic 124, the AEC logic 126, and the adder logic 128 in the audio processing device 120 of Figures 2A to 2C . However, it should be noted that various implementations and embodiments can use various and possibly different components to perform the operations of the method in Figures 2A to 2C . For example, in various embodiments, various semiconductor devices - such as SoC, field programmable gate array (FPGA), programmable logic device (PLD), application specific integrated circuit (ASIC), or other integrated circuit devices - can be configured with firmware instructions that are operable to perform the operations of the method in Figures 2A to 2C when executed by a processor and / or other hardware components (e.g., microcontroller, state machine, etc.). In another example, in various embodiments, an IC device can include a single-chip audio controller or a multi-chip audio controller configured to perform the operations of the method in Figures 2A to 2CThe description of the method as performed by the spatial filtering logic, AEC logic, and adder logic in the audio processing device should be considered illustrative rather than restrictive.
[0074] Figure 2A illustrates a method for nonlinear echo cancellation that can be implemented in a system with a single AEC logic (e.g., such as Figure 1A system 100A in ). In Figure 2A , according to input operation 202, the reference signal x and the multi-channel microphone digital signal are provided as inputs to the spatial filtering logic in the audio processing device having the spatial filter f SF . For example, the reference signal x provided in other ways is continuously provided to the spatial filtering logic and the AEC logic. The multi-channel microphone digital signal is a digital multi-channel signal generated based on audio signals from multiple microphones in a microphone array acoustically adjacent to the speaker. Therefore, the multi-channel microphone digital signal includes both linear echo and nonlinear echo picked up by the microphones in the microphone array. As part of operation 202, one of the microphone signals in the multi-channel microphone signal (e.g., the i-th) microphone signal is also provided as an input to the adder logic in the audio processing device. The i-th microphone signal y i can be predetermined (e.g., based on the known / fixed arrangement of the speaker relative to the microphone array), or can be randomly selected from the channels of the multi-channel microphone signal during operation.
[0075] In operation 204, based on the reference signal x, the spatial filter f in the spatial filtering logic SF is applied to the multi-channel microphone signal and a spatially filtered signal is generated (e.g., according to equation (6) above). The generated spatially filtered signal carries both the linear echo and the nonlinear echo included in the i-th signal y i . Then the spatially filtered signal is provided as an input to the linear AEC logic in the audio processing device.
[0076] In operation 206, the AEC logic adaptively calculates the coefficients w of its linear adaptive filter. The AEC logic applies the linear adaptive filter together with its coefficients w to the spatially filtered signal (e.g., according to equation (7) and equation (8) above) to generate a cancellation signal cancellation signal Estimate both the linear echo signal and the non-linear echo signal included in the i-th microphone signal y i and then provide the cancellation signal as an input to the adder logic. Additionally, in some embodiments, the AEC logic may be configured to utilize both the reference signal x and the spatially filtered signal . For example, the AEC logic may be configured to use the reference signal x for DTD. The AEC logic may also be configured to use the spatially filtered signal for its background filter and the reference signal x for its foreground filter, and select one (e.g., “best”) output from the output of the background filter and the output of the foreground filter to minimize the cancellation of the near-end speech during a two-way call scenario.
[0077] In operation 208, the adder logic receives the cancellation signal and the i-th microphone signal y i . The adder logic performs a digital summation based on the cancellation signal and based on the i-th microphone signal y i and generates an output signal e (e.g., according to equation (8) above). In fact, the output signal e approximates the target speech signal s captured by the i-th microphone (e.g., ). In this way, both the linear echo signal and the non-linear echo signal are cancelled from the output signal e.
[0078] In operation 210, the output signal e is provided as an output (e.g., provided to a host application). Additionally, the output signal e may also be provided as feedback to the AEC logic, which adaptively calculates the linear adaptive coefficients w using the output signal e.
[0079] Figure 2B illustrates a method for non-linear echo cancellation that may be implemented in a system (e.g., such as system 100B in Figure 1B ) having multiple instances of AEC logic - where one AEC instance is applied to each microphone signal / channel. In Figure 2B , according to input operation 212, the reference signal x and the multi-channel microphone digital signal are provided as inputs to the spatial filtering logic having a spatial filter f SF in an audio processing device. For example, the reference signal x provided in some other way for transmission to the speaker is continuously provided to the spatial filtering logic. The multi-channel microphone digital signal is a digital multi-channel signal generated based on audio signals from multiple microphones in a microphone array acoustically adjacent to the speaker. Thus, the multi-channel microphone digital signal Includes both linear echo and non-linear echo picked up by microphones in a microphone array. As part of operation 212, a reference signal x is also provided to one or more instances of the AEC logic, and the multi-channel microphone signal is also provided as an input to the adder logic of the audio processing device.
[0080] In operation 214a, a spatial filter f in the spatial filtering logic is applied to the multi-channel microphone signal SF based on the reference signal x and a spatially filtered signal is generated (e.g., according to equation (6) above). The generated spatially filtered signal carries both the linear echo and the non-linear echo included in the multi-channel signal . Then the spatially filtered signal is provided as an input to each instance of the multiple instances of the linear AEC logic of the audio processing device.
[0081] In operation 216, each instance of the AEC logic adaptively calculates the coefficients w of its corresponding linear adaptive filter i . Each instance of the AEC logic applies its linear adaptive filter with its corresponding coefficients w i to the spatially filtered signal (e.g., according to equation (10) above) to generate a cancellation signal Thus, the cancellation signal is a multi-channel echo estimation signal that estimates both the linear echo signal and the non-linear echo signal included in all microphone signals y of the multi-channel microphone signal. Then the multi-channel cancellation signal i is provided as an input to the adder logic. Additionally, in some embodiments, one or more instances of the AEC logic may be configured to utilize both the reference signal x and the spatially filtered signal . For example, one or more instances of the AEC logic may be configured to use the reference signal x for DTD. Each instance of the AEC logic may also be configured to use the spatially filtered signal for its background filter and the reference signal x for its foreground filter, and select one (e.g., "best") output from the output of the background filter and the output of the foreground filter to minimize the cancellation of the near-end speech during a two-way conversation scenario. In operation 218, the adder logic receives the multi-channel cancellation signal
[0082] and the multi-channel microphone signal and The adder logic is based on the multi-channel cancellation signal and is based on the multi-channel microphone signal to perform digital summation and generate a multi-channel output signal (e.g., according to equation (10) above). The multi-channel output signal is provided as an input to the spatial filter f in the spatial filtering logic SF for operation 214b.
[0083] In operation 214b, the spatial filter f in the spatial filtering logic SF is applied to the multi-channel output signal (e.g., with appropriate filter coefficients) to generate a spatially filtered output signal (e.g., according to equation (9) above). In various embodiments, the spatial filter f in operation 214b SF may also be configured to receive the reference signal x, the multi-channel microphone signal and / or the multi-channel cancellation signal from one or more of these signals, and use any and / or all of these signals when generating the multi-channel output signal . In fact, the output signal e approximates the target speech signal s captured by the microphones in the microphone array (e.g., ). In this way, both the linear echo signal and the non-linear echo signal are eliminated from the output signal e.
[0084] In operation 220, the output signal e is then provided as an output (e.g., provided to a host application). Additionally, the multi-channel output signal may also be provided as feedback to each instance of the AEC logic, and each instance of the AEC logic uses the multi-channel output signal to adaptively calculate its corresponding linear adaptive coefficients w of its corresponding linear adaptive filter i .
[0085] Figure 2C shows a method for non-linear echo cancellation that can be implemented in a system (e.g., such as the system 100C in Figure 1C ) having a single AEC logic applied to the spatial filter output. In Figure 2C , according to input operation 222, the reference signal x and the multi-channel microphone digital signal are provided as inputs to the audio processing device having the spatial filter f SFSpatial filtering logic. For example, a reference signal x provided otherwise for transmission to a speaker is continuously provided to the spatial filtering logic. The reference signal x is also provided as an input to the AEC logic of the audio processing device. The multi-channel microphone digital signal is a digital multi-channel signal generated based on audio signals from a plurality of microphones in a microphone array acoustically adjacent to the speaker. Thus, the multi-channel microphone digital signal includes both linear echoes and non-linear echoes picked up by the microphones in the microphone array.
[0086] In operation 224, a spatial filter f in the spatial filtering logic is applied to the multi-channel microphone signal SF based on the reference signal x and a spatially filtered signal is generated (e.g., according to equation (6) above). The generated spatially filtered signal carries both the linear echo and the non-linear echo included in the multi-channel signal . Also as part of operation 224, the same or different instances of the spatial filter f in the spatial filtering logic are applied to the multi-channel microphone signal SF to generate a spatially filtered microphone signal (e.g., according to equation (11) above). The spatially filtered microphone signal is generated using filter coefficients for the microphones in the microphone array and thus carries the spatially amplified speech estimation signal and the attenuated echo estimation signal both included in the multi-channel signal (e.g., ). After generation, the spatially filtered signal is provided as an input to the linear AEC logic, and the spatially filtered microphone signal is provided as an input to the adder logic of the audio processing device.
[0087] In operation 226, the AEC logic adaptively calculates the coefficients w of its linear adaptive filter. The AEC logic applies the linear adaptive filter with its coefficients w to the spatially filtered signal (e.g., according to equation (12) above) to generate a cancellation signal The cancellation signal estimates both the linear echo signal and the non-linear echo signal included in the spatially filtered microphone signal . Then the cancellation signal Provided as input to the adder logic. Additionally, in some embodiments, the AEC logic may be configured to utilize both the reference signal x and the spatially filtered signal For example, the AEC logic may be configured to use the reference signal x for DTD. The AEC logic may also be configured to use the spatially filtered signal For its background filter and the reference signal x for its foreground filter, and select one (e.g., "best") output from the output of the background filter and the output of the foreground filter to minimize the cancellation of the near-end speech during a two-way call scenario.
[0088] In operation 228, the adder logic receives the cancellation signal And the spatially filtered microphone signal The adder logic performs a digital summation based on the cancellation signal And the spatially filtered microphone signal And generates an output signal e (e.g., according to equation (12) above). In fact, the output signal e approximates the target speech signal s captured by the microphones in the microphone array (e.g., ). In this way, both the linear echo signal and the nonlinear echo signal are eliminated from the output signal e, and the target speech signal (e.g., in the case of a two-way call) is avoided from being eliminated from the output signal e.
[0089] In operation 230, the output signal e is then provided as an output (e.g., provided to a host application). Additionally, the output signal e may also be provided as feedback to the AEC logic, which adaptively calculates the linear adaptive coefficient w using the output signal e.
[0090] The techniques described herein provide significant improvements, making it possible to apply nonlinear echo cancellation to embedded systems, edge devices, and other systems with limited computational capabilities. For example, conventional nonlinear echo cancellation methods typically result in solutions that are computationally too expensive (e.g., Volterra filters, Hammerstein filters, neural networks, etc.) or not robust enough to account for the time-varying nature of the nonlinearity (e.g., preprocessing filters). In contrast, the techniques described herein provide a practical, robust solution for eliminating nonlinear echo using a linear filter, which is both robust and computationally suitable for systems / devices with limited computational capabilities.
[0091] Figures 3A to 3BA figure showing a simulation study performed to verify the effectiveness of the proposed solution in accordance with techniques for non - linear echo cancellation described herein. Generally, such simulation studies are a reliable mechanism for predicting signal processing results and are often used as the first step in building practical solutions in the field of digital signal processing. Establish Figure 3A The specific simulation study reflected in Figure 3A simulates a system with 6 circular microphones uniformly arranged with a radius of 3 cm. Non - linearities in the simulated system are modeled using second - order polynomial approximation and third - order polynomial approximation commonly found in consumer speakers. The linear impulse response of the simulated system is modeled using 85 delay - line taps, which are set to operate at 15 kHz to simulate multiple echoes from the same source signal. The linear adaptive filter of the system is under - modeled by 20% to simulate real - world conditions.
[0092] Figure 3A Figure 300 shows a graph of the average error magnitude of three different echo - cancellation mechanisms. Specifically, line 304 shows the error - magnitude results of AEC using a conventional linear adaptive filter without non - linear echo cancellation. Line 306 shows the error - magnitude results of AEC using a linear adaptive filter for non - linear echo cancellation in accordance with the techniques described herein. Line 308 shows the error - magnitude results of AEC using a non - linear Volterra filter for non - linear echo cancellation of "known" non - linearities. As Figure 3A shown, non - linear echo cancellation in accordance with the techniques described herein (line 306) has almost the same convergence as AEC with a linear filter without non - linear echo cancellation (line 304), but provides an additional 10 dB of additional cancellation when compared to AEC with a linear filter (line 304). At the same time, non - linear echo cancellation in accordance with the techniques described herein (line 306) has echo - cancellation performance that is substantially equivalent to AEC with a non - linear Volterra filter for "known" linearities (line 308).
[0093] Figure 3B Figure 310 shows a graph of the modeled linear response (line 316) using the techniques for non - linear echo cancellation described herein relative to the modeled linear response (line 314) of a conventional method using AEC with a linear filter and the ideal response (line 312). As Figure 3B shown, the non - linear echo - cancellation mechanism in accordance with the techniques described herein (line 316) is able to model acoustic coupling better than the conventional method (line 314), while achieving results comparable to the ideal echo cancellation (line 312) of the simulated system.
[0094] Figure 3A and Figure 3BThe simulation results herein show that the techniques described herein for nonlinear echo cancellation have almost the same convergence characteristics as conventional AECs using linear filters, but provide an additional 10 dB of echo cancellation over conventional methods and have a nonlinear echo cancellation performance substantially equivalent to that of an AEC with a nonlinear Volterra filter for "known" nonlinearities.
[0095] The techniques described herein for nonlinear echo cancellation are applicable to systems using multiple microphones. In various embodiments, the techniques described provide for generating a spatially filtered signal to estimate a nonlinear reference signal by spatial filtering of a multi-channel microphone signal, which is provided to an AEC with a linear adaptive filter for echo cancellation. Compared to conventional methods using nonlinear filters or preprocessing filters, the techniques described herein provide several benefits. For example, the solutions according to the techniques described herein provide low complexity, which reduces the computational cost of echo cancellation and makes such solutions practical for devices with limited computational capabilities (e.g., SoCs and IoT devices). Additionally, the solutions according to the techniques described herein are more robust because they are able to track changes in the nonlinearity over time and improve the linear adaptive filter estimates by reducing the statistical bias due to the nonlinearity.
[0096] In various embodiments, the techniques described herein for nonlinear echo cancellation can be applied to smart speakers and IoT edge devices and can be implemented in firmware and / or hardware depending on the availability of local device resources. A smart speaker is a multimedia device with a built-in speaker and microphone that enables human-machine interaction via voice commands. An IoT edge device is an entry point for IoT cloud-based services. For example, in a smart speaker embodiment with multiple microphones, the techniques described herein can significantly save computational cycles while providing "good enough" performance not only after a BF direction change but also providing fast convergence for all other types of echo path changes while maintaining noise robustness. In an IoT edge device embodiment, the techniques described herein can enhance the voice signals received by the IoT edge device for a backend system that may be running automatic speech recognition.
[0097] The techniques described herein for nonlinear acoustic echo cancellation can be implemented on various types of audio processing devices. Figure 4 An example audio processing device configured according to the techniques described herein is shown. In Figure 4In the embodiment shown, the audio processing device 400 may be a single-chip IC device fabricated on a semiconductor die or a single-chip IC fabricated as an SoC. In other embodiments, the audio processing device 400 may be a multi-chip module encapsulated in a single semiconductor package or multiple semiconductor packages arranged or mounted on a common substrate such as a PCB. Thus, Figure 4 the audio processing device 400 in
[0098] should be considered in an illustrative rather than a limiting sense. In addition to other components, the audio processing device 400 includes audio I / O logic 410, a DSP 420, a CPU 432, a read-only memory (ROM) 434, a random access memory (RAM) 436, and a host interface 438. The DSP 420, CPU 432, ROM 434, RAM 436, and host interface 438 are coupled to one or more buses 430. The DSP 420 is also coupled to the audio I / O logic 410 via a multi-channel bus. The audio I / O logic 410 is coupled to the speaker-microphone assembly 110.
[0099] The speaker - microphone assembly 110 includes one or more speakers 112 and a microphone array 114. The microphone array 114 includes a plurality of microphones that are arranged to detect sound waves from a desired sound source (e.g., human speech), but can also detect / record sound waves from an undesired sound source (e.g., an echo such as from the speaker 112). The speaker 112 is coupled to a digital - to - analog converter (DAC) circuitry in the audio I / O logic 410. The speaker 112 is configured to receive an analog audio signal from the DAC circuitry and emit the audio signal as sound waves. The microphone array 114 is coupled to an analog - to - digital converter (ADC) circuitry in the audio I / O logic 410. The microphone array 114 is configured to receive sound waves from various sound sources and convert them into an analog audio signal that is sent to the ADC circuitry. In some embodiments, some or all of the microphones in the microphone array 114 can share the same communication channel to the ADC circuitry in the audio I / O logic 410 via a suitable multiplexer and buffer. In other embodiments, each microphone in the microphone array 114 can have a separate communication channel to the ADC circuitry in the audio I / O logic 410 and a separate instance of the ADC circuitry. In some embodiments (e.g., a smart phone), the speaker 112 and the microphone array 114 can be integrally formed as the same speaker - microphone assembly 110. In some embodiments (e.g., a conference phone device), the speaker 112 and the microphone array 114 can be separate components arranged on a common substrate (e.g., a PCB) that is mounted inside or on the housing of the speaker - microphone assembly 110. In yet another embodiment, the speaker - microphone assembly 110 can not have a housing, but can be formed by the acoustic proximity of the speaker 112 and the microphone array 114.
[0100] The audio I / O logic 410 includes various logic blocks and circuitry configured to process signals transmitted between the DSP 420 and the speaker - microphone assembly 110. For example, the audio I / O logic 410 includes DAC circuitry and ADC circuitry. The DAC circuitry includes a DAC, an amplifier, and other circuitry suitable for signal processing (e.g., circuitry for input matching, amplitude limiting, compression, gain control, parametric or adaptive equalization, phase shift, etc.) that is configured to receive a modulated digital signal from the DSP 420 and convert it into an analog audio signal for the speaker 112. The ADC circuitry includes an ADC, an amplifier, and other circuitry suitable for signal processing (e.g., circuitry for input matching, amplitude limiting, compression, gain control, parametric or adaptive equalization, phase shift, etc.) that is configured to receive an analog audio signal from the microphones in the microphone array 114 and convert it into a modulated digital signal that is sent to the DSP 420.
[0101] The DSP 420 includes various logic blocks and circuitry configured to process digital signals transferred between the audio I / O logic 410 and various components coupled to the bus 430. For example, the DSP 420 includes circuitry configured to receive digital audio data (e.g., a series of bytes) from other components in the audio processing apparatus 400 and convert the received audio data into a modulated digital signal (e.g., a bitstream) that is sent to the audio I / O logic 410. The DSP 420 also includes circuitry configured to receive the modulated digital signal from the audio I / O logic 410 and convert the received signal into digital audio data. In Figure 4 the illustrated embodiment, the DSP 420 includes a Break-In Sub-System (BISS) logic 422. The BISS logic 422 includes a spatial filtering logic block (with a spatial filter f SF ) configured according to the non-linear echo cancellation techniques described herein, an AEC logic block with a linear adaptive filter, and an adder logic block. The spatial filtering logic block may implement a BF logic block or may be implemented as part of a BF logic block. The BISS logic 422 also includes a control register configured to control the operation of the spatial filtering logic block, the AEC logic block, and the adder logic block, and a shared memory (e.g., RAM) that shares signal data within its logic block and with other blocks of the DSP 420 and / or with various components in the audio processing apparatus 400. The BISS logic 422 may also include a Programmable State Machine (PSM). The PSM may be implemented as a microcode engine that includes its own microcontroller, which can extract instructions from a microcode memory and use the shared memory to obtain operands for its instructions. The PSM is configured to exercise fine-grained control over the hardware circuitry by programming internal hardware registers (IHRs) co-located with the hardware functions it controls.
[0102] The bus 430 may include one or more buses, such as a system interconnect and a peripheral interconnect. The system interconnect may be a single-level or multi-level Advanced High-Performance Bus (AHB) configured as an interface to couple the CPU 432 to other components of the audio processing apparatus 400 and as a data and control interface between various components and the peripheral interconnect. The peripheral interconnect may be an Advanced eXtensible Interface (AXI) bus that provides a primary data and control interface between the CPU 432 and its peripherals and other resources (e.g., system resources, I / O blocks, Direct Memory Access (DMA) controllers, etc.), which can be programmed to transfer data between peripheral blocks without burdening the CPU.
[0103] The CPU 432 includes one or more processing cores configured to execute instructions that may be stored in the ROM 434, RAM 436, or flash memory (not shown). The ROM 434 is a read-only memory (or other suitable non-volatile storage medium) configured to store boot routines, configuration parameters, and other firmware parameters and settings. The RAM 436 is a volatile memory configured to store data and firmware instructions accessed by the CPU 432. The flash memory, if present, may be an embedded or external non-volatile memory (e.g., NAND flash, NOR flash, etc.) configured to store data, programs, and / or other firmware instructions.
[0104] The host interface 438 may include control registers, data registers, and other circuitry configured to transfer data between the DSP 420 and a host (not shown). The host may be an on-chip microcontroller subsystem, an off-chip IC device (e.g., an SoC), and / or an external computer system. The host may include its own CPU operable to execute host applications or other firmware / software configured to (among other functions) send, receive, and / or process audio data. In some embodiments, multiple communication circuits and / or hosts may be instantiated on the same audio processing device 400 to provide communication via various protocols (e.g., such as Bluetooth and / or Wi-Fi) for audio signals and / or other signals sent, received, or otherwise processed by the audio processing device 400. In some embodiments (e.g., such as a smart phone), an application processor (AP) may be instantiated as an on-chip host coupled to the host interface 438 to provide execution of various applications and software programs.
[0105] In operation, the DSP 420 receives audio data (e.g., a series of bytes) via the bus 430 (e.g., from the host interface 438). The DSP 420 converts the received audio data into a modulated digital signal (e.g., a bit stream), which is sent as a reference signal x(n) to the BISS logic 422. The modulated digital signal is also sent to the audio I / O logic 410. The audio I / O logic 410 converts the received digital signal into an analog audio signal that is sent to the speaker 112. The microphones in the microphone array 114 pick up sound waves from proximal speech as well as linear and non-linear echoes (if any) from the speaker 112. The microphones in the microphone array 114 convert the received sound waves into corresponding analog audio signals that are sent to the audio I / O logic 410. The audio I / O logic 410 converts the received analog audio signals into a multi-channel microphone digital signal that is sent to the BISS logic 422 in the DSP 420
[0106] In some embodiments, the audio processing apparatus 400 may be configured with a single AEC logic (e.g., in system 100A in Figure 1A ) to perform the method for non-linear echo cancellation shown in Figure 2A . In some embodiments, the audio processing apparatus 400 may be configured with multiple instances of AEC logic, where one AEC instance is applied to each microphone signal / channel (e.g., in system 100B in Figure 1B ) to perform the method for non-linear echo cancellation shown in Figure 2B . In some embodiments, the audio processing apparatus 400 may be configured with a single AEC logic that is applied to the spatial filter output (e.g., in system 100C in Figure 1C ) to perform the method for non-linear echo cancellation shown in Figure 2C . It should be noted that the audio processing apparatus 400 may be configured in a system with other components and hardware circuits, and for this reason, the description of the audio processing apparatus implemented in the operating environments of systems 100A to 100C in Figures 1A to 1C should be considered illustrative rather than restrictive.
[0107] Figure 5 is a block diagram showing a host apparatus 500 according to various embodiments. The host apparatus 500 may fully or partially include and / or operate the host 140 in FIG. 1, and / or be coupled to the audio processing apparatus 400 through the host interface 438 to Figure 4 . Figure 5 The host apparatus 500 shown in
[0108] may operate as a stand-alone device or may be connected (e.g., networked) to other machines. In a networked deployment, the host apparatus 500 may be implemented as a server blade in a cloud-based physical infrastructure, a server or client machine in a server-client network, a peer machine in a P2P (or distributed) network, etc.The host device 500 can be implemented in various form factors (e.g., an on-chip device, a computer system, etc.). Multiple sets of instructions can be executed within the host device 500 to cause the host device 500 to perform one or more of the operations and functions described herein. For example, in various embodiments, the host device 500 can be a SoC device, an IoT device, a server computer, a server blade, a client computer, a personal computer (PC), a tablet, a set-top box (STB), a personal digital assistant (PDA), a smart phone, a network device, a speakerphone, a handheld multimedia device, a handheld video player, a handheld gaming device, or any other machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. When the host device 500 is implemented as an on-chip device (e.g., a SoC, an IoT device, etc.), the components shown therein can reside on a common carrier substrate such as an IC die substrate, a multi-chip module substrate, etc. When the host device 500 is implemented as a computer system (e.g., a server blade, a server computer, a PC, etc.), the components shown therein can be separate integrated circuits and / or discrete components arranged on one or more PCB substrates. Additionally, although only a single host device 500 is shown in Figure 5 , in various operating environments, the term "device" can also generally be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the operations and functions described herein.
[0109] The host device 500 includes a processor 502, a memory 503, a data storage interface 504, a display interface 505, a communication interface 506, a user input interface 507, and an audio interface 508 coupled to one or more buses 501. When the host device 500 is implemented as an on-chip device, the bus 501 can include one or more on-chip buses such as a system interconnect (e.g., single-level or multi-level AHB) and a peripheral interconnect (e.g., AXI bus). When the host device 500 is implemented as a computer system, the bus 501 can include one or more computer buses such as a chipset north / south bridge (which coordinates communication between the processor 502 and other components) and various peripheral buses (e.g., PCI, serial ATA, etc. that coordinate communication to various computer peripherals).
[0110] The host device 500 includes a processor 502. When the host device 500 is implemented as an on-chip device, the processor 502 can include an ARM processor, a RISC processor, a microprocessor, an application processor, a controller, a dedicated processor, a DSP, an ASIC, an FPGA, etc. When the host device 500 is implemented as a computer system, the processor 502 can include one or more CPUs.
[0111] The host device 500 also includes a memory 503. The memory 503 may include a non-volatile memory (e.g., ROM) for storing static data and instructions for the processor 502, a volatile memory (e.g., RAM) for storing data and executable instructions for the processor 502, and / or a flash memory for storing firmware (e.g., control algorithms) that can be executed by the processor 502 to implement at least a portion of the operations and functions described herein. Portions of the memory 503 may also be dynamically allocated to provide caching, buffering, and / or other memory-based functions. The memory 503 may also include a removable memory device that can store one or more sets of software instructions. Such software instructions may also be sent or received over a network via the communication interface 506. During the execution of software instructions by the host device 500, the software instructions may also reside entirely or at least partially on a non-transitory computer-readable storage medium and / or within the processor 502.
[0112] The host device 500 also includes a data storage interface 504. The data storage interface 504 is configured to connect the host device 500 to a storage device configured for persistent storage of data and information used by the host device 500. Such data storage devices may include persistent storage media of various media types, including but not limited to magnetic disks (e.g., hard disks), optical storage disks (e.g., CD-ROMs), magneto-optical storage disks, solid state drives, universal serial bus (USB) flash drives, and the like.
[0113] The host device 500 also includes a display interface 505 and a communication interface 506. The display interface 505 is configured to connect the host device 500 to a display device (e.g., a liquid crystal display (LCD), a touch screen, a computer monitor, a television screen, etc.) and provide software and hardware support for a display interface protocol. The communication interface 506 is configured to send data to and receive data from other computing systems / devices. For example, the communication interface 506 may include a USB controller and bus for communicating with USB peripherals, a network interface card (NIC) for communicating over a wired communication network, and / or a wireless network card that can implement various wireless data transfer protocols such as IEEE 802.11 (Wi-Fi) and Bluetooth.
[0114] The host device 500 further includes a user input interface 507 and an audio interface 508. The user input interface 507 is configured to connect the host device 500 to various input devices, such as alphanumeric input devices (e.g., a touch-sensitive keyboard or a typewriter-style keyboard), a pointing device that provides spatial input data (e.g., a computer mouse), and / or any other suitable human-machine interface device (HID) that can transmit user commands and other user-generated information to the processor 502. The audio interface 508 is configured to connect the host device 500 to various audio devices (e.g., a microphone, a speaker, etc.) and provide software and hardware support for various audio input / output.
[0115] Various embodiments of the techniques for non-linear acoustic echo cancellation described herein may include various operations. These operations may be performed and / or controlled by hardware components, digital hardware, and / or firmware and / or combinations thereof. As used herein, the term "coupled to" may mean directly connected or indirectly connected through one or more intermediate components. Any of the signals provided through various on-chip buses may be time-division multiplexed with other signals and may be provided through one or more common on-chip buses. Additionally, the interconnection between circuit components or blocks may be shown as a bus or shown as a single signal line. Each of the buses in the bus may alternatively be one or more single signal lines, and each of the single signal lines in the single signal line may alternatively be a bus.
[0116] Certain embodiments may be implemented as a computer program product, which may include instructions stored on a non-transitory computer-readable medium such as volatile memory and / or non-volatile memory. These instructions may be used to program and / or configure one or more devices including a processor (e.g., a CPU) or its equivalent (e.g., such as a processing core, a processing engine, a microcontroller, etc.) such that when executed by the processor or its equivalent, these instructions cause the device to perform the described operations for non-linear echo cancellation. The computer-readable medium may also include one or more mechanisms for storing or transmitting information in a form readable by a machine (e.g., such as a device or a computer) (e.g., software, a processing application, etc.). The non-transitory computer-readable storage medium may include, but is not limited to, electromagnetic storage media (e.g., a floppy disk, a hard disk, etc.), optical storage media (e.g., a CD-ROM), magneto-optical storage media, read-only memory (ROM), random access memory (RAM), erasable programmable memory (e.g., EPROM and EEPROM), flash memory, or other non-transitory media now known or later developed suitable for storing information.
[0117] Although the operations of the circuits and blocks are shown and described herein in a particular order, in some embodiments, the order of operations of each circuit / block can be changed such that certain operations can be performed in the reverse order, or such that certain operations can be performed at least partially concurrently and / or in parallel with other operations. In other embodiments, the instructions or sub-operations of different operations can be performed in an intermittent and / or alternating manner.
[0118] In the foregoing specification, the invention has been described with reference to specific exemplary embodiments of the invention. However, it will be apparent that various modifications and changes can be made to the invention without departing from the broader spirit and scope of the invention as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. An audio processing system, comprising: A speaker, which is configured to receive a reference signal; A microphone array, which is configured to provide a multi-channel microphone signal, the multi-channel microphone signal including both linear echo and non-linear echo; Spatial filtering logic, which is configured to receive the reference signal and the multi-channel microphone signal, and generate a spatially filtered signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; Acoustic echo canceller logic, which is at least configured to: Receive the spatially filtered signal; and Apply a linear adaptive filter using the spatially filtered signal to generate a cancellation signal, the cancellation signal estimating both the linear echo and the non-linear echo of the multi-channel microphone signal; and A logic block, which is configured to receive the cancellation signal and generate an output signal based at least on the cancellation signal.
2. The audio processing system according to claim 1, wherein The system further includes beamformer logic, and the beamformer logic includes the spatial filtering logic.
3. The audio processing system according to claim 1, wherein The acoustic echo canceller logic is configured to: periodically calculate the filter coefficients of the linear adaptive filter based on the output signal.
4. The audio processing system according to claim 1, wherein, The logic block is configured to: generate the output signal based on the cancellation signal and based on the microphone signal from one channel in the multi-channel microphone signal.
5. The audio processing system according to claim 1, further comprising a plurality of instances of the acoustic echo canceller logic and at least two instances of the spatial filtering logic.
6. The audio processing system according to claim 5, wherein: The plurality of instances of the acoustic echo canceller logic are configured to: generate the cancellation signal as a multi-channel echo estimation signal; and The logic block is further configured to: Generate a multi-channel output signal based on the multi-channel echo estimation signal and the multi-channel microphone signal; and Apply an instance of the spatial filtering logic using the multi-channel output signal to generate the output signal.
7. The audio processing system according to claim 1, wherein: The spatial filtering logic is further configured to: generate a spatially filtered microphone signal based on the multi-channel microphone signal; and The logic block is configured to: generate the output signal based on the cancellation signal and the spatially filtered microphone signal.
8. The audio processing system according to claim 1, further comprising a host, the host being configured to: receive the output signal from the logic block and perform speech recognition.
9. The audio processing system according to claim 8, wherein, The host is configured to: Generate the reference signal; and Provide the reference signal to the speaker and the spatial filtering logic.
10. The audio processing system according to claim 8, wherein, The spatial filtering logic, the acoustic echo canceller logic, and the logic block are arranged on a semiconductor device coupled to the host through a network.
11. The audio processing system according to claim 1, wherein, The system is one of a speakerphone, a smart speaker, and a smart phone.
12. A semiconductor device for audio processing, the semiconductor device including a digital signal processor, the digital signal processor being configured to: Receive a reference signal sent to a speaker; Receiving a multi-channel microphone signal from a microphone array, wherein, The multi-channel microphone signal includes both linear echo and non-linear echo; Applying a spatial filter to the reference signal and the multi-channel microphone signal to generate a spatially filtered signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; Applying a linear adaptive filter to the spatially filtered signal to generate a cancellation signal, the cancellation signal estimating both the linear echo and the non-linear echo of the multi-channel microphone signal; and Generating an output signal based at least on the cancellation signal.
13. The semiconductor device according to claim 12, wherein, The digital signal processor is configured to generate the output signal based on the cancellation signal and based on a microphone signal from one channel of the multi-channel microphone signal.
14. The semiconductor device according to claim 12, wherein, The digital signal processor includes a plurality of instances of acoustic echo canceller logic having a linear adaptive filter, and wherein: The plurality of instances of acoustic echo canceller logic are configured to: generate the cancellation signal as a multi-channel echo estimation signal; and The digital signal processor is further configured to: Generate a multi-channel output signal based on the multi-channel echo estimation signal and the multi-channel microphone signal; and Apply the spatial filter to the multi-channel output signal to generate the output signal.
15. The semiconductor device according to claim 12, wherein, The digital signal processor is further configured to: Apply the spatial filter to the multi-channel microphone signal to generate a spatially filtered microphone signal; and Generate the output signal based on the cancellation signal and the spatially filtered microphone signal.
16. The semiconductor device according to claim 12, wherein, The digital signal processor includes: beamformer logic including the spatial filter; acoustic echo canceller logic including the linear adaptive filter; and a logic block including adder logic configured to generate the output signal.
17. A method for acoustic echo cancellation, the method comprising: Receiving a reference signal transmitted to a speaker; Receiving a multi-channel microphone signal from a microphone array acoustically adjacent to the speaker, wherein the multi-channel microphone signal includes both linear echo and non-linear echo; Generating, by a processing device, a spatially filtered signal by applying a spatial filter to the reference signal and the multi-channel microphone signal, wherein the spatially filtered signal carries both the linear echo and the non-linear echo of the multi-channel microphone signal; Generating, by the processing device, a cancellation signal by applying a linear adaptive filter to the spatially filtered signal, wherein the cancellation signal estimates both the linear echo and the non-linear echo of the multi-channel microphone signal; and Generating, by the processing device, an output signal based at least on the cancellation signal.
18. The method according to claim 17, wherein, Generating the output signal includes: using the cancellation signal and a microphone signal from one channel of the multi-channel microphone signal.
19. The method according to claim 17, wherein: Generating the cancellation signal includes: using a plurality of linear adaptive filters to generate the cancellation signal as a multi-channel echo estimation signal; and Generating the output signal further includes: Generate a multi-channel output signal based on the multi-channel echo estimation signal and the multi-channel microphone signal; and Apply the spatial filter using the multi-channel output signal to generate the output signal.
20. The method according to claim 19, wherein: Generating the cancellation signal includes: applying the spatial filter using the multi-channel microphone signal to generate a spatially filtered microphone signal; and Generating the output signal further includes: generating the output signal based on the cancellation signal and the spatially filtered microphone signal.
Citation Information
Patent Citations
Audio signal processing
CN102968999A
Conferencing Apparatus that combines a Beamforming Microphone Array with an Acoustic Echo Canceller
US20170134849A1