Method and apparatus for variable pitch echo cancellation
By employing a variable step size adaptive filtering method based on the BLUE standard in acoustic echo cancellation, the convergence and robustness issues of adaptive filters under dual-talk conditions are solved, achieving uniform convergence and fast acoustic path tracking in the frequency domain.
Patent Information
- Application Number
- CN202180070673.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-15
- Filing Date
- 2021-09-27
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2041-09-27
AI Technical Summary
Existing acoustic echo cancellation technology tends to diverge or converge slowly in dual-talk situations, making it difficult to find a balance between maintaining convergence and quickly tracking changes in the acoustic path.
A variable step size adaptive filtering method based on the BLUE standard is adopted. By estimating the power spectral density of the useful signal and the echo-to-signal energy ratio, the update step size of the adaptive filter is dynamically adjusted to ensure uniform convergence in the frequency domain and maintain robustness in dual-talk situations.
Effective echo cancellation is achieved in dual-talk scenarios, ensuring rapid convergence and stability of the adaptive filter under acoustic path changes, reducing complexity and improving the robustness of echo cancellation.
Smart Images

Figure CN116420315B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present description relates to a method and a device for echo cancellation. BACKGROUND
[0002] In the context of simultaneous sound capture and playback, it is appropriate to use a process involving acoustic echo cancellation (or AEC hereafter).
[0003] As Figure 1 illustrated, a device item comprises at least one loudspeaker HP and at least one microphone MIC capturing a microphone signal y(t). The loudspeaker HP is fed with a signal x(t) which, when emitted by the loudspeaker HP, is transformed by the environment (possible reverberation, Larsen effect, or others) and captured by the microphone together with a useful signal s(t) currently acquired by the microphone. The microphone signal y(t) is thus composed of:
[0004] - a useful signal s(t) (possibly involving speech signal data from a conversation, a voice command, or others), also called "signal of interest s(t)" or "local signal s" hereafter, depending on the context, and
[0005] - an echo signal z(t) emitted by a sound playback system comprised in the device item and composed of one or more loudspeakers HP.
[0006] This echo signal is associated with a direct path between the microphone and the playback system, as well as any reflections of the signal x(t) in the propagation environment.
[0007] The whole acoustic path can be modeled by a finite impulse response filter w whose length depends on the characteristics of the propagation environment, so that:
[0008] z(t) = x(t) * w(t)
[0009] The operation of removing the contribution of the echo signal z(t) from the microphone signal y(t) is called "acoustic echo cancellation" (or AEC). The process performing this operation can include deriving the echo signal from an estimate of the acoustic path. This operation is called "adaptive filtering". By subtracting the estimated echo signal from the microphone signal y(t), an estimated useful signal is derived, as illustrated below:
[0010]
[0011] Adaptive filtering is usually performed based on the correlation between the microphone signal and the loudspeaker signal, exploiting the statistical independence between the signal x(t) emitted by the loudspeaker and the signal s(t) of interest. In practice, in order to track the variations of the acoustic path represented by the filter w (and for convenience, referred to hereinafter as acoustic path w), it is appropriate to perform this processing over a short-term horizon. These variations can typically be manifested when the person speaking moves through the room forming the environment.
[0012] As a consequence of this short-term processing, the statistical independence between the signal s(t) of interest and the loudspeaker signal x(t) can no longer be maintained in certain cases, except for the trivial case where the signal s(t) is zero. Indeed, this independence no longer holds when computed over a short time window of a few tens to a few hundreds of milliseconds, typically corresponding to the conventional frame length of digital signals.
[0013] In these so-called "double talk" situations, i.e. when the useful signal s(t) is non-zero, the consequence is a bias in the estimate of the acoustic path, degrading the echo cancellation. Less complex solutions based on a processing using stochastic gradients, such as the "Normalized Least Mean Square" (NLMS) technique and its derivatives, are very sensitive to the presence of the local signal s(t). In these double talk situations, if the filter continues to adapt, it can even diverge, ultimately leading to an echo amplification, which is contrary to the desired effect. Furthermore, in order to be efficient, adaptive filtering solutions must be robust to double talk situations, while being able to quickly track the variations of the acoustic path.
[0014] Ideally, this filtering should only process the data in play, i.e. the reference signal x(t) and the microphone signal y(t).
[0015] In order to overcome the double talk situations, certain known adaptive filtering processing solutions implement a double talk detection (DTD) system. For example, a system of this type is described in the reference [@jung2005new], the disclosure details of which are given in the appendix at the end of the present specification. Such a system disables the adaptation during the periods identified as double talk. However, in practice, the DTD is subject to a detection delay, which can lead to an echo. On the other hand, in this particular case of binary decision, the adaptation of the filter is frozen during the double talk periods, which in practice can be distracting if the filter has not yet completed its convergence, resulting in a perceptible residual echo.
[0016] Instead, other methods have been proposed to derive an adaptive step-size in the estimation of the acoustic path. In the known references, this step-size is continuous. Unlike binary decision schemes such as DTD, such implementations make it possible to continue to track the acoustic path, including during double talk. These types of adaptation are usually derived per frequency band, as follows:
[0017]
[0018] where ΔW is the update at each time instant k and each frequency f of the estimated acoustic channel .
[0019] On the one hand, working at frequencies makes it possible to converge more uniformly over the entire frequency range considered. On the other hand, the spectral sparsity of the signal makes it possible to continue to estimate the acoustic channel in one frequency band, while freezing in another. Certain methods, called "variable step-size" or VSS, propose to adjust the adaptation ΔW according to different criteria.
[0020] It has been attempted to smooth the random adaptation by freezing iterations considered too random, in particular to avoid random updates due to the presence of double talk.
[0021] It has also been attempted to directly measure the local speech presence rate, in the form of the ratio between the energy of the local signal and the energy of the echo signal , but this adaptation becomes fixed when this ratio is too high. Since the estimated value of the variance is particularly large in noise, they are directly used to adjust the adaptive step-size making these schemes ineffective in practice: they freeze the adaptation too much, slowing down the speed of convergence, or insufficiently limit the mismatch during double talk.
[0022] Other methods are based on an optimal solution for the adaptive step-size, which guarantees the minimum variance of the estimated filter in the case of minimum echo. This criterion is called "BLUE" for "Best Linear Unbiased Estimator" [@trump1998frequency]. Updating the acoustic path ΔW in the adaptive filtering process according to this criterion (k) allows to limit the residual echo related to the variations of the adaptive filter around its solution (minimum variance). However, in practice, the BLUE expression depends on the second-order statistics of the signal s(t) (more precisely, on its statistical autocorrelation matrix Γ s ), which is unknown and generally varies over time, as is the case with non-stationary signals such as speech in general [@van2007double}. Thus, the solution proposed in [@trump1998frequency] is not fully satisfactory. SUMMARY
[0023] The invention improves this situation.
[0024] A method is proposed for processing a signal y(t) from at least one microphone of a device item further comprising at least one loudspeaker intended to be fed with a signal x(t),
[0025] The processing of the signal y(t) from the microphone:
[0026] - aims at least at limiting the echo effect caused by a microphone capturing the sound emitted by a loudspeaker in the environment of the device item, said sound emitted by the loudspeaker and any possible acoustic reflections thereof along an acoustic path w(t) from the loudspeaker to the microphone,
[0027] - and comprises, in order to limit this echo effect, determining a useful signal s(t) by subtracting from the signal y(t) from the microphone an estimate x(t)* of the echo signal given by applying to the signal x(t) fed to the loudspeaker a filter which can be adapted with variable step size so as to take into account the variations of the acoustic path w(t) over time,
[0028] The method wherein:
[0029] * the signal x(t) fed to the loudspeaker is obtained in a successive form of signal samples over time frames, and
[0030] * the adaptive filter is produced at each frame k of samples as a function of an update AW (k) of the acoustic path w to this frame k and by applying a normalization Λ satisfying criteria selected to be minimum variance, said normalization Λ being a function of parameters representative of statistical expectations of the useful signal s(t).
[0031] Such an implementation provides an acoustic echo cancellation solution which is particularly robust to double talk situations, as described below.
[0032] In one embodiment, the above mentioned selected criteria are of the "Best Linear Unbiased Estimate" "BLUE" type.
[0033] In the case of a matrix representation of the useful signal s, said statistical expectations can be written as E{ss H}(s H the conjugate transpose of the designated matrix s). For example, in the time domain, in the case of a simple scalar representation, it can depend on a time parameter t, and can be written E{s(t)s(t- t)}.
[0034] In the frequency domain, the statistical expectation can be represented by a parameter corresponding to the power spectral density. Thus, in an implementation produced by an adaptive filter, for example, in the domain of frequency subbands f, its expression can be E{s(f)s(f- f)} = Γ s (f) a function of the corresponding parameter. In particular, the normalization Λ(f) represented in the frequency domain is itself a function of the power spectral density Γ s of the useful signal s(f).
[0035] In such an embodiment, the normalization Λ (k) is more precisely defined as a function of the power spectral density Γ of the useful signal s, and is also defined as a function of the power spectral density Γ of the signal x supplied to the loudspeaker.
[0036] In this embodiment, in a matrix representation in which f represents the row index (and here also the frequency subband index) and b represents the column index, the normalization Λ (k) (f,b) can be given by the following formula:
[0037] where μ ∈ [0, 2[ and where γ is a chosen positive coefficient (in the context of a practical implementation, this choice can be empiric).
[0038] The power spectral density Γ of the useful signal s can itself be estimated as a function of the power spectral density Γ of the signal y captured by the microphone and of a representation of the echo-to-signal energy ratio .
[0039] In this embodiment, in a matrix representation in which f represents the row index and b represents the column index, the power spectral density Γ of the useful signal s is given by the following formula:
[0040]
[0041] where A is a chosen positive limit (for example, in practice a chosen positive term "very large" such as 10 10 ), and is the power spectral density of the useful signal s evaluated for the previous frame k-1, in the frequency subband f and for the partition b.
[0042] The representation of the echo-to-signal energy ratio Itself can be estimated as a function of the inter-power spectral density between the signal y from the microphone and the signal X intended for the loudspeaker .
[0043] For example, in a matrix representation where f denotes the row index and b denotes the column index, the representation of the echo-to-signal energy ratio can be given by:
[0044] where β is a positive forgetting factor smaller than 1, the symbol (k-1) denotes the expression determined for the previous frame (k-1).
[0045] In this expression, the inter-power spectral density can be given by:
[0046]
[0047] where {α, δ, η, ξ} ∈ ]0, 1].
[0048] The power spectral density of the signal intended for the loudspeaker X and the signal y from the microphone, in a matrix representation where X is a matrix and y is a vector, can be given by:
[0049]
[0050]
[0051] where α and η are forgetting factors greater than 0 and smaller than 1. Here, the squared norm of a matrix (or vector), denoted as |. |, is defined as the matrix 2 of the squared norms of each element of the matrix
[0052] In embodiments providing an advantage for the estimation of the adaptive filter, the latter can be represented by successive partitions. Thus, in such embodiments, the filter w can be of the finite impulse response type and have a length of N samples. In particular, it is subdivided into partitions w b each having L samples.
[0053] In such embodiments, one can estimate the matrix This matrix corresponds to the representation of the partitions w b in the transform domain (for example, in the aforementioned frequency subbands), such that W = [w1,..., w B ], and represents the filter in the transform domain, where w b = Fw b , M > L, where F is a field transform matrix.
[0054] It will be noted that in this embodiment, the column index "b" can correspond to a partition index w b However, the matrix representation given above with row index f and column index b can be applied to cases other than the case involving a partition of the filter. As a direct illustrative example, the formula given above remains valid in a degraded embodiment for example with b = 1, thus not involving a partition.
[0055] Moreover, for each time frame of M samples of the signal x(t) intended to be supplied to the loudspeaker, denoted , the transform of the signal intended to be supplied to the loudspeaker and corresponding to the last B frames x b forms the matrix such that X = [x1,..., x B ], where x b = Fx b . For the time frame of the signal y(t) coming from the microphone the final vector
[0056] This vector y can be constructed such that:
[0057]
[0058] In this format, the update of the acoustic path ΔW (k) for the current frame k is given by where:
[0059] - denotes the Hadamard product,
[0060] - is a matrix given by one of the following:
[0061] G = FF H and G = I M ,
[0062] - is a matrix representing the above normalization, and
[0063] - e (k) is the a priori error estimated from the signals x and y according to the frame k.
[0064] The a priori error is given by:
[0065]
[0066] In which the acoustic path ΔW(k) An embodiment of updating the function h from the current frame k to the next frame k+1, updating the adaptive filter, can estimate this update for the current frame k, and the update on the acoustic path is given by a relation of the type: (k+1) = W (k) + ΔW (k) .
[0067] The specification also relates to a computer program comprising instructions for implementing the above method when the program is executed by a processor.
[0068] It also relates to a device for processing a signal y(t) coming from at least one microphone, comprising a processor configured to perform the method as defined above. BRIEF DESCRIPTION OF DRAWINGS
[0069] Other features, details and advantages will become clear on reading the following detailed description and on analyzing the attached drawings, in which:
[0070] Figure 1 An item of equipment is shown in which the purposes of the present specification can be implemented, according to one embodiment.
[0071] Figure 2 A processing is shown in order to deliver the useful signal mentioned above, according to one embodiment.
[0072] Figure 3 A processing is shown in order to deliver the update on the estimate of the acoustic path mentioned above, according to one embodiment.
[0073] Figure 4 An apparatus for implementing the purposes of the present specification is shown, according to one embodiment. DETAILED DESCRIPTION
[0074] The following drawings and description contain some elements that are essentially determinative. Thus, they not only contribute to a better understanding of the present disclosure, but also, when applied, they contribute to its definition.
[0075] The present specification proposes hereinafter an acoustic echo cancellation solution that is robust to double talk situations. It is based on a processing involving an adaptive filtering, such as a NLMS processing, usually applied continuously to each frame of successive frames. A frame is understood here to mean a given number of successive samples of the signal x(t) fed to the loudspeaker, which is of course considered to be digital.
[0076] In one embodiment, the filter used for the adaptive filtering is partitioned (each partition can or can not correspond to the length of a frame), preferably in the frequency domain (technique referred to here as "Partitioned Block Frequency Domain NLMS" or "PBFD-NLMS"). This type of technique is for example described in the reference [@borrallo1992implementation].
[0077] More specifically here, the solution is based on the derivation of the BLUE optimal step size, but without adding auxiliary information, directly estimating the necessary statistics from the reference and microphone signals. This makes it possible to calculate ΔW (k) , which can be the case in the references of the prior art, in particular [@gil2014frequency].
[0078] Such an embodiment guarantees, without auxiliary information other than that directly derived by the processing itself, a convergence close to optimal in terms of convergence speed, zero bias at convergence, and absence of divergence in the double talk situation.
[0079] When expressed in the frequency domain, the adaptive filtering makes it particularly possible to control and normalize the update of the acoustic path independently of the frequency bands involved. Thus, in addition to reducing the complexity, the solution also benefits from a more homogeneous convergence over the entire frequency range considered.
[0080] It also makes it possible, via the operation of the partitions, associated with the filtering in the frequency domain, to estimate at each iteration of the processing the time-frequency representation W (k) of the acoustic path w. This makes it possible to implement different adaptive strategies depending on the partitions. It also makes it possible to guarantee better convergence in the case of very long filters.
[0081] Such a processing allows to derive a step size that optimizes both the behavior in double talk situation and the acoustic path tracking.
[0082] Figure 2 The different steps of the adaptive filtering solution are illustrated. At each iteration of the adaptive filtering processing, a frame of L new samples of the signals x(t) and y(t) is considered and produces L new samples. In step S1, it is determined whether it is necessary to initialize the acoustic path to be considered (for example, at the beginning of a conversation between the loudspeaker and the other party), in which case the initialization of the acoustic path is performed in step S2. Otherwise, in step S3, the acoustic echo cancellation AEC processing is directly started. In step S4, a time frame of the reference signal x(t) is retrieved, and in the described example, a projection is applied on it in the frequency domain (for example, in the domain of frequency subbands) in step S5, to obtain a frequency representation x (k) . A similar processing is performed (step S6) on each time frame of the microphone signal y(t) to obtain a projection y (k) in the frequency domain in step S7. Based on the frames x(t) and y(t) (or, as described herein, on their frequency representations), an echo cancellation processing is applied in step S8, in order to estimate a priori error e (k) .
[0083] Embodiments of adaptive filtering for echo cancellation are described below by generating filters by partitioning according to the "partitioned block" technique.
[0084] Considering the target acoustic path modeled by the finite impulse response filter w(t) which is N samples long w, and is subdivided into partitions We can estimate the matrix This matrix corresponds to the frequency transform of the partitions w b such that
[0085] W = [w1,..., w B ], where w b = Fw b , M > L.
[0086] F is a domain transform matrix, for example here a redundant Discrete Fourier Transform (DFT) matrix, such that each element is characterized by: In practice, the redundancy is achieved by padding with zeros in the time domain.
[0087] In the same way, we consider a time frame of the reference signal x(t) comprising M samples, and we form a matrix corresponding to the frequency transform of the last B frames x b such that:
[0088] X = [x1,..., x B ], where x b = Fx b .
[0089] By further considering the time frame of the microphone signal y(t) We can represent vectors Make:
[0090]
[0091] To avoid the problems associated with convolution operations performed in the frequency domain, this processing is based on overlap-preservation operations (OLS). Here, exponent... (k) This reflects the k-th iteration of the process. After initializing W... (0) X (0) y (0) After considering other characteristics, the prior error can be calculated. Let's continue processing:
[0092]
[0093] in This represents the Hadama product, and (.) * Represents the conjugate of a matrix or vector.
[0094] As mentioned above, the redundancy of the DFT is achieved by means of zero padding, which is found in the expression of the prior error, and this advantageously makes it possible to avoid artifacts caused by circular convolution.
[0095] The method then continues to calculate the update of the acoustic path in step S9. as follows:
[0096] in
[0097] In the embodiment where updates are optimal (due to zero-padding), we set G = FF. H (This is called a "constraint" update).
[0098] In an embodiment where the update is suboptimal, we can instead set G=I. M (The update is called "unconstrained"). Such an implementation has the advantage of consuming fewer resources.
[0099] Step S10 then aims to calculate W (k+1) In order to update the acoustic path:
[0100] E (k+1) =W (k) +ΔW (k) .
[0101] And in step S11, when using W(k) x (k) After the convolution, the useful signal is obtained and returned to the time domain.
[0102] Figure 3 The steps to calculate the update of the acoustic path W (k) are detailed, in particular the optimal normalisation term Λ that allows the robustness inherent to the double-talk situation.
[0103] In order for the echo cancellation solution to be robust in the above situations, the spectral normalisation term Λ (k) that satisfies the BLUE criterion is chosen. This can be achieved thanks to the knowledge of the power spectral density (PSD) of the microphone signal x(t) and of the local signal s(t).
[0104] In the reference [@van 2007 double], the BLUE is obtained at the cost of strong assumptions on the local signal that must accept an autoregressive model, be considered as speech and use an error prediction method. On the other hand, [@trump 1998 frequency] achieves the BLUE by estimating the PSD of the error signal e (k) in post, at the cost of less stability and operating in the case of strong constraints on the local signal (stationary coloured noise).
[0105] The proposed solution explained below overcomes the constraints used above thanks to the robust estimation of the DSP over time.
[0106] Assumptions:
[0107] - The power spectral density (PSD) estimate of X (respectively of y) for each frequency and each bin The spectral inter-
[0108] relation between the microphone signal and the reference signal and its power spectral inter-density
[0109] The power spectral density of the local signal s for each frequency and each bin is designated as
[0110] Finally, we represent the matrix as (ESR stands for "echo signal ratio") that represents the ratio of the energy of the echo and of the local signal for each frequency band and each bin. After the initialisation and other characteristics, the process performs the following operations:
[0111] - Estimate of the Power Spectral Density (PSD):
[0112] where
[0113] where
[0114]
[0115] where {a, δ, η, ξ} e ]0, 1].
[0116] - Estimate of the Instantaneous Echo Signal Ratio (ESR):
[0117]
[0118] where β e ]0, 1].
[0119] - Estimate of the Normalization Λ (k)
[0120]
[0121] Then, by applying a method within the meaning of the specification, the normalization parameter is expressed, in order to satisfy the BLUE criterion, as:
[0122] where μ e [0, 2[.
[0123] This last expression of the normalization parameter involves the term which is a function of the estimate of the echo signal ratio for each frame k, which can eventually be the only parameter to be estimated within the meaning of the specification (of course with the signal x(t)).
[0124] It should be noted that, otherwise, by applying the teachings of the prior art described for example in [@borrallo1992implementation], the classical technique of PBFD-NLMS would have the normalization parameter of the filter represented as follows:
[0125] where μ e [0, 2[ without involving any measure for estimating the echo signal ratio.
[0126] Now, by taking Figure 3 The procedure is similar to the one described in the previous section, the first step S20 starts with a test to determine whether the power spectral density estimates are initialized or not. If this is the case, the individual spectral densities of the microphone signal y and the reference signal x are initialized in step S21. Otherwise, the procedure for estimating the spectral normalization factor Λ is started directly in step S22. In step S23, the current frequency frame of the microphone signal is retrieved and in step S24 the current frequency frame of the reference signal is retrieved in order to estimate the inter-spectral density in step S25. Then, the power spectral density of the microphone signal is estimated in step S26 and the power spectral density of the reference signal is estimated in step S27 in order to derive the estimate of the instantaneous echo signal ratio (ESR) from it in step S28 as described above. The estimate of the spectral normalization factor Λ is derived from it in step S29, from which the update of the acoustic path can be determined in step S30: (k)
[0127] The BLUE criterion is fulfilled by the adaptive filtering performed in the time domain typically amounts to finding the solution of the equation such that:
[0128]
[0129] where is the microphone signal vector, is the matrix of the loudspeaker signals, and is the autocorrelation matrix of the signal s.
[0130] Current echo cancellation methods are based on adaptive filtering in the frequency domain. The method presented in [@trump1998frequency} proposes to fulfill a normalized version of the BLUE criterion by finding the channel solution in the frequency domain, such that:
[0131]
[0132] where Γ s (resp. Γ x ) is the diagonal matrix of the power spectral density of the signal s (resp. x).
[0133] However, the local signal s is unknown. In practice, it is difficult to obtain an estimator satisfying the BLUE criterion if no other information or model on s is available.
[0134] If the local signal s and the reference signal x are de-correlated, i.e. E{Xs T} = 0 (where E{·} denotes the expectation operator), the solution described in [@borrallo1992implementation} can only produce an unbiased estimate of the channel (Thus satisfying the BLUE criterion). In practice, this condition is only satisfied when the signal s is white noise. In all other cases, to achieve unbiased estimation... It is necessary to normalize the factor Λ (k) Add a small portion of the local signal variance s to the denominator, i.e., E{ss} H} or its power spectral density (PSD) in the frequency domain (taking parameters in the above equations) (in the form of).
[0135] Because of the expression E{ss H Since Γ is unknown, the proposed solution overcomes the constraints used above due to the robust estimation of the power spectral density over time, particularly the occurrence of Γ in the denominator described above. s The power spectral density of the local signal s.
[0136] The processing described above can be particularly useful in situations where sound needs to be captured and played back simultaneously. The most common use cases are hands-free phones (where a person speaking at a distance hears their own delayed voice—an echo—mixed with the other party's voice), interactions with voice assistants (where responses from the dialogue system and / or music played on the voice assistant are mixed with the user's commands and interfere with speech recognition), walkie-talkies, video conferencing systems, etc.
[0137] The apparatus for implementing the above method is in Figure 4 The text indicates that it can also be generated by... Figure 1 The two modules on the left (applied to adaptive filtering and subtraction of the microphone-captured signal y(t)) will be used for illustration. (See reference...) Figure 4 The device typically includes a first input interface IN1 for receiving a signal y(t) acquired from a microphone MIC, and a second input interface IN2 for receiving a signal (e.g., a telecommunications signal, such as a voice or music signal) to be played back on a speaker HP. The device includes a processor PROC capable of cooperating with a memory MEM to process the audio signal and deliver a signal x(t) intended for supplying the speaker HP via a first output interface OUT1 included in the device. Specifically, the memory MEM stores at least instruction data of a computer program according to one aspect of this specification. In an exemplary embodiment, the instruction data can be read by the processor PROC to perform the aforementioned processing and, in particular, applied to the signal y(t) from the microphone to deliver a useful signal s(t) via a second output interface OUT2 included in the device.
[0138] Of course, this is an exemplary embodiment, wherein here generally the useful signal s(t) can be transmitted via the output interface OUT2 to e.g. a remote party. In this case, for example, the interface OUT2 can be connected to a communication antenna or a router of a telecommunication network NET. The same is true for the input interface IN2, which receives a signal to be played through the loudspeaker from "outside".
[0139] For example, in a device such as a voice assistant, it is also possible that a double talk situation is present when the user speaks a voice command while the assistant is responding to e.g. a previous command. In this case, at least part of the response of the voice assistant can be issued locally from the content of the memory MEM, e.g. without necessarily using a remote server and a telecommunication network. Furthermore, the useful signal s(t) can be interpreted locally by the processor PROC only in order to respond to the voice command from the user. Therefore, the interfaces IN2 and OUT2 can not be necessary.
[0140] A typical use case for the processing referred to as voice assistant "double talk" processing includes, for example, listening to music through the loudspeaker of the voice assistant while the user is speaking a command to wake up the assistant (wake-up word). In this case, in order to be able to correctly detect the actually spoken command signal s(t), it is suggested to cancel the played music x(t)*w(t) from the sound signal y(t) captured in the environment in which the assistant has just been placed (with its reverberation).
[0141] Of course, the present invention is not limited to the above-described embodiments and can be extended to other variants.
[0142] For example, above, an actual embodiment was described in which the signals are processed in the domain of frequency subbands. However, the present invention can also be implemented in the time domain by utilizing a parameter such as the expectation E{ss T} which is equivalent to the power spectral density Γ s of the local signal s in the frequency domain.
[0143] Therefore, the estimation of the power spectral density described above can be effectively performed in one possible but not necessary implementation.
[0144] The same is true for the adaptive filter partitioning, which can also work in the subband domain, but this is not necessary either.
[0145] Furthermore, in Figure 4In a central aspect, a compact device item has been shown, which comprises an echo cancellation device (which can thus be shown by a processor PROC, a memory MEM, and at least one input interface and at least one output interface), as well as a microphone MIC and a loudspeaker HP. In variant embodiments, the device on the one hand and the one or more microphones and the one or more loudspeakers on the other hand can be located at different sites, connected by e.g. a telecommunication network or a local area network (powered by a home gateway) or other means.
[0146] Appendix: References
[0147] [@borrallo1992implementation]: Borrallo, J. P., & Otero, M. G. (1992). On the implementation of a partitioned block frequency domain adaptive filter (PBFDAF) for long acoustic echo cancellation. Signal Processing, 27(3), 301-315.
[0148] [@trump1998frequency]: Trump, T. (1998, May). A frequency domain adaptive algorithm for colored measurement noise environment. In Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP’98 (Cat. No. 98CH36181) (Vol. 3, pp. 1705-1708). IEEE.
[0149] [@jung2005new]: Jung, H. K., Kim, N. S., & Kim, T. (2005). A new double-talk detector using echo path estimation. Speech communication, 45(1), 41-48.
[0150] [@van2007double]: Van Waterschoot, T., Rombouts, G., Verhoeve, P., & Moonen, M. (2007). Double-talk-robust prediction error identification algorithms for acoustic echo cancellation. IEEE Transactions on Signal Processing, 55(3), 846-858.
[0151] [@gil2014frequency]: Gil-Cacho, J. M., Van Waterschoot, T., Moonen, M., & Jensen, S. H. (2014). A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 22(12), 2074-2086.
Claims
1. A method for processing a signal y(t) from at least one microphone of a device item, said device item further comprising at least one loudspeaker intended to be fed with a signal x(t), the processing of said signal y(t) from the microphone: including, in order to limit the echo effect, determining an estimate of the echo signal y(t) by subtracting from the signal y(t) coming from the microphone the estimate of the echo signal y(t) given by the useful signal s(t), said filter may be adapted with variable step size in order to take into account the variations over time of the acoustic path w(t) from the loudspeaker to the microphone, said method wherein: * said signal x(t) fed to the loudspeaker is obtained in a continuous form of signal samples over time frames, and * Adaptive filter is generated at each frame k of the samples as a function of the update ΔW (k) of the acoustic path w(t) of this frame k and by applying a normalization Λ satisfying criteria selected to be minimum variance, said normalization A is a function of parameters representative of statistical expectations of the useful signal s(t), wherein said adaptive filter is produced in a domain f of frequency subbands, and said normalization A is a function of: - the power spectral density of the useful signal s and - the power spectral density of the signal x supplied to the loudspeaker a power spectral density of the signal y captured by the microphone and a representation of the echo to signal energy ratio to estimate a power spectral density of the wanted signal s 2. The method according to claim 1, wherein the chosen criterion is of the "BLUE" type for "Best Linear Unbiased Estimator".
3. The method of claim 1, wherein, In a matrix representation where f denotes the row index and b denotes the column index, the normalization Λ (k) (f,b) is given by wherein μ ∈ [0, 2[ and wherein γ is a chosen positive coefficient.
4. The method of claim 1, wherein, In a matrix representation where f denotes the row index and b denotes the column index, the power spectral density of the useful signal s is given by where A is a selected positive limit, and is the power spectral density of the useful signal estimated for the previous frame k-1, in the frequency sub-band f and for the partition b.
5. The method of claim 1, wherein, Representation of echo-to-signal energy ratio estimated as a function of the cross-power spectral density between the signal y, assumed to come at least from the microphone, and the signal x, intended to be supplied to the loudspeaker. 6. The method of claim 5, wherein, In a matrix representation where f denotes the row index and b denotes the column index, the representation of the echo-to-signal energy ratio is given by: where β is a positive forgetting factor smaller than 1, the symbol (k-1) denotes an expression determined for a previous frame (k-1).
7. The method of claim 6, wherein the inter-power spectrum density is given by: wherein {a, δ, η, ξ} ∈ ]0, 1].
8. The method according to claim 6, wherein the power spectral densities of: - the signal intended to be fed to the loudspeaker, represented by the matrix X, and - the signal from the microphone, represented by the vector y, are respectively given by: and wherein a and η are forgetting factors greater than 0 and less than 1.
9. The method of claim 1, wherein the adaptive filter is a finite impulse response filter w of N samples long and is subdivided into P partitions w b each having L samples. 10. The method of claim 9, wherein the estimation matrix corresponds to an expression in a transform domain of the partitions w b such that W = [w1,..., w B ], and represents a filter in the transform domain, where w b = Fw b , M > L, where F is a domain transform matrix, And wherein, For each time frame of M samples of the signal x(t) intended for the loudspeaker, denoted as the transform of the last B frames x b forms a matrix such that X = [x1,..., x B ], where x b = Fx b , and For a time frame of the signal y(t) from the microphone Forming a vector 11. The method according to claim 10, wherein said vector y is such that:
12. The method of claim 10, wherein, an update AW of the acoustic path for the current frame k (k) is given by wherein: denotes a Hadamard product, - is a matrix given by one of the following formulae: G = FF H and G = I M , - is a matrix representing the above normalization, and e (k) is the a priori error estimated from the signals x and y of frame k.
13. The method according to claim 12, wherein said a priori error is given by:
14. The method of claim 1, wherein, According to the following type of relationship: W (k+1) = W (k) + ΔW (k) , as a function of the update ΔW (k) of the estimate of the acoustic path for the current frame k, the adaptive filter is updated from the current frame k to the next frame k+1.
15. A non-transitory computer storage medium storing instructions of a computer program which, when executed by a processor, cause the implementation of the method according to any one of claims 1 to 14.
16. An apparatus for processing a signal y(t) from at least one microphone and comprising a processor configured to perform the method according to any one of claims 1 to 14.