Method and device for variable pitch echo cancellation
The method uses a variable step size adaptive filter with spectral normalization to address echo cancellation challenges during double talk, ensuring robust and efficient echo removal across frequency bands.
Patent Information
- Application Number
- JP2023523159
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-15
- Filing Date
- 2021-09-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2041-09-27
AI Technical Summary
Existing echo cancellation methods struggle to effectively remove echoes in situations of simultaneous sound capture and playback, particularly during 'double talk' scenarios, leading to biased estimates and potential echo amplification due to the loss of statistical independence between the signal of interest and the loudspeaker signal.
A method employing a variable step size adaptive filter that normalizes the acoustic path update based on the power spectral density of the useful signal and echo-to-signal energy ratio, ensuring robustness in double-talk situations and optimal convergence without additional information.
The solution provides near-optimal convergence and prevents echo amplification, maintaining effective echo cancellation even during double talk, with reduced complexity and improved frequency domain uniformity.
Smart Images

Figure 0007775309000132 
Figure 0007775309000133 
Figure 0007775309000134
Abstract
Description
[Technical Field]
[0001] The present specification relates to methods and devices for echo cancellation. [Background technology]
[0002] In situations of simultaneous sound capture and playback, it is appropriate to use processing that includes acoustic echo cancellation (or hereinafter "AEC").
[0003] As shown in Figure 1, the equipment item comprises at least one loudspeaker HP and at least one microphone MIC capturing a microphone signal y(t). The loudspeaker HP is supplied with a signal x(t), which, when emitted by the loudspeaker HP, is transformed by the environment (possible reverberation, Larsen effect, etc.) and captured by the microphone together with a useful signal s(t) currently collected by the microphone MIC. The microphone signal y(t) is therefore a useful signal s(t) (possibly relating to speech signal data from a conversation, voice commands, etc.), also referred to below as the "signal of interest s(t)" or "local signal s" depending on the context, and - an echo signal z(t) emitted by a sound reproduction system contained in an equipment item and consisting of one or more loudspeakers HP; It consists of:
[0004] This echo signal is associated with the direct path between the microphone and the playback system as well as with any reflections of the signal x(t) in the propagation environment.
[0005] The entire acoustic path can be modeled by a finite impulse response filter w, whose length depends on the properties of the propagation environment, as follows: z(t)=x(t)*w(t)
[0006] The operation that consists of removing the contribution of the echo signal z(t) from the microphone signal y(t) is called "acoustic echo cancellation" (or AEC).
[0007]
number
[0008] The acoustic path
[0009]
number
[0010] This operation is called "adaptive filtering".
[0011]
number
[0012] is the echo signal estimated from the microphone signal y(t) as follows:
[0013]
number
[0014] It is derived by subtracting
[0015]
number
[0016] The adaptive filtering is generally performed based on the correlation between the microphone signal and the loudspeaker signal, exploiting the statistical independence between the signal x(t) emitted by the loudspeaker and the signal of interest s(t). In practice, it is appropriate to perform this process with a short deadline in order to track changes in the acoustic channel represented by the filter w (hereafter referred to as the acoustic path w for convenience). These changes may generally appear when the speaker is moving around the room forming the environment.
[0017] The consequence of this short-term processing is that the statistical independence between the signal of interest s(t) and the loudspeaker signal x(t) no longer holds in some situations, except in the insignificant case where the signal s(t) is zero. In fact, this independence no longer holds when calculated over short time periods of tens to hundreds of milliseconds, which typically correspond to conventional frame lengths of digital signals.
[0018] In these situations, known as "double talk," i.e., when the useful signal s(t) is nonzero, the result is a biased estimate of the acoustic channel, degrading echo cancellation. Less complex solutions, such as those based on processes using stochastic gradients, such as the "normalized least mean squares" (NLMS) technique and its derivatives, are highly sensitive to the presence of the local signal s(t). If the filter continues to adapt during these double talk situations, it may even diverge, ultimately causing echo amplification, the opposite of the desired effect. To be effective, adaptive filtering solutions must also be robust to double talk situations while being able to quickly track changes in the acoustic path.
[0019] Ideally, this filtering should only process the data being played back, ie the reference signal x(t) and the microphone signal y(t).
[0020] To overcome double-talk situations, some known adaptive filtering solutions implement double-talk detection (DTD) systems. This type of system is described, for example, in reference [@jung2005new], the disclosure of which is provided in detail in the appendix at the end of this specification. Such systems disable adaptation during periods identified as double-talk. However, DTDs in practice suffer from detection delays, which can lead to echoes. On the other hand, in this particular case of binary decisions, filter adaptation is frozen during double-talk periods, which can actually be distracting and result in perceptible residual echoes if the filter has not yet finished converging.
[0021] Other methods instead propose deriving an adaptive step size in the acoustic path estimation. In known references, this step size is continuous. Such implementations, unlike binary decision methods such as DTD, allow for continued tracking of the acoustic path, including during double talk periods. These types of adaptations are usually derived by frequency band as follows:
[0022]
number
[0023] where ΔW is the estimated acoustic channel
[0024]
number
[0025] at each instant k and each frequency f.
[0026] On the one hand, working in frequency allows for a more uniform convergence over the entire frequency range considered. On the other hand, the spectral sparsity of the signal allows for continuing to estimate the acoustic channel in one frequency band while freezing the estimate in another. Several methods, called "variable step size" or VSS, propose adjusting the adaptation ΔW according to different criteria.
[0027] Attempts have been made to smooth the probability adaptation by freezing iterations that are deemed too random, in particular to avoid random updates due to the presence of double talk.
[0028] Local signal energy
[0029]
number
[0030] and that of the echo signal
[0031]
number
[0032] Attempts have also been made to measure the local speech presence rate directly, in the form of a ratio between σ and σ, but this adaptation has stalled when this ratio is too high.
[0033]
number
[0034] Since estimates of ,are particularly noisy, using them directly to adjust the adaptation step size would actually render these techniques ineffective; they would either freeze the adaptation too much, slowing down convergence, or they would not sufficiently bound the mismatch during double-talk periods.
[0035] Another method is based on the optimal solution of the adaptation step size, which guarantees the minimum variance of the estimated filter, given the minimum echo. This criterion is called "BLUE" for "Best Linear Unbiased Estimate" [@trump1998frequency]. According to this criterion, the adaptive filtering process optimizes the acoustic path ΔW (k) Updating Γ(t) makes it possible to limit the residual echo associated with the variation of the adaptive filter around its solution (minimum of variance). In practice, however, the BLUE expression is determined by the second-order statistics of the signal s(t) (and more precisely its statistical autocorrelation matrix Γ), which are not only unknown but also generally time-varying, as is the case with non-stationary signals such as speech. s The solution presented in @trump1998frequency is therefore still not entirely satisfactory. [Prior art documents] [Non-patent literature]
[0036] [Non-Patent Document 1] [@Borrallo1992implementation]: Borrallo, JP and Otero, MG (1992). On the implementation of a partitioned block frequency domain adaptive filter (PBFDAF) for long acoustic echo cancellation. Signal Processing, 27(3), 301~315 [Non-patent document 2] [@Trump1998frequency]: Trump, T.(1998, May). A frequency domain adaptive algorithm for colored measurement noise environment. In Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP'98(Cat.No.98CH36181) (Vol.3, pp. 1705~1708) [Non-patent document 3] [@Jung2005new]: Jung, HK, Kim, NS, and Kim, T.(2005). A new double-talk detector using echo path estimation. Speech communication, 45(1), 41-48 [Non-patent document 4] [@van2007double]: Van Waterschoot, T., Rombouts, G., Verhoeve, P., and Moonen, M. (2007). Double-talk-robust prediction error identification algorithms for acoustic echo cancellation. IEEE Transactions on Signal Processing, 55(3), 846–858 [Non-patent document 5] [@gil2014frequency]: Gil-Cacho, JM, Van Waterschoot, T., Moonen, M., and Jensen, SH (2014). A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 22(12), 2074~2086 Summary of the Invention [Means for solving the problem]
[0037] The present invention improves this situation.
[0038] A method is proposed for processing a signal y(t) coming from at least one microphone of an item of equipment, the item of equipment further comprising at least one loudspeaker intended to be supplied with a signal x(t), The processing of the above signal y(t) from the microphone is - aiming at least to limit the echo effect induced by the microphone capturing the sound emitted by the loudspeaker in the environment of the equipment item, said sound emitted by the loudspeaker and any possible acoustic reflections following an acoustic path w(t) from the loudspeaker to the microphone, - Applying a filter to the signal x(t) fed to the loudspeaker to limit echo effects
[0039]
number
[0040] The echo signal given by applying
[0041]
number
[0042] Determination of the useful signal s(t) by subtracting the estimate of from the signal y(t) coming from the microphone
[0043]
number
[0044] Contains filters
[0045]
number
[0046] is adaptable with a variable step size to take into account the changes over time of the acoustic path w(t) above, the signal x(t) fed to the loudspeaker is obtained in the form of a time-dependent succession of frames of signal samples, * Adaptive filter
[0047]
number
[0048] is the update ΔW of the acoustic path w(t) for each frame k of samples. (k) The method is performed by applying a normalization Λ that satisfies a criterion chosen for minimum variance at this frame k, where the normalization Λ is a function of a parameter that represents the statistical expectation of the useful signal s(t).
[0049] As will be described in more detail below, such an implementation presents an acoustic echo cancellation solution that is particularly robust in double-talk situations.
[0050] In one embodiment, the criterion chosen, as mentioned above, is the "BLUE" type of "Best Linear Unbiased Estimator."
[0051] The above statistical expectation is given by E{ss H} can be written as (s H denotes the conjugate transpose of a matrix s). For example, in the time domain, and for expressions that are simply scalars, it may depend on a time parameter τ and can be written as E{s(t)s(t-τ)}.
[0052] In the frequency domain, the above statistical expectation can be expressed by parameters corresponding to the power spectral density. Thus, in an implementation in which the adaptive filter is generated, for example, in the domain of the frequency subband f, its formula is the power spectral density Γ of the useful signal s(f) s In particular, the normalized Λ(f) above, expressed in the frequency domain, can be a function of the parameters corresponding to the power spectral density Γ of the useful signal s. s is a function of the parameters corresponding to
[0053] In such an embodiment, the normalized Λ (k) is the power spectral density of the useful signal s
[0054]
number
[0055] as a function of , and the power spectral density of the signal x fed to the loudspeaker
[0056]
number
[0057] is more precisely defined as a function of
[0058] In this embodiment, in a matrix representation where f denotes the row index (and here the index of the frequency subband) and b denotes the column index, the normalized Λ (k) (f,b) can be given by:
[0059]
number
[0060] , μ∈[0,2[, where γ is a chosen positive coefficient (this choice can be empirical in light of the actual implementation).
[0061] power spectral density of the useful signal s
[0062]
number
[0063] is itself the power spectral density of the signal y captured by the microphone
[0064]
number
[0065] as a function of and the expression of the echo-to-signal energy ratio
[0066]
number
[0067] can be estimated as a function of
[0068] In this embodiment, the power spectral density of the useful signal s is expressed in matrix representation, where f denotes the row index and b denotes the column index:
[0069]
number
[0070] is given by:
[0071]
number
[0072] where A is the chosen positive limit (e.g., 10 10 (In practice, "extremely large" selected positive terms such as
[0073]
number
[0074] is the power spectral density of the useful signal s evaluated for the previous frame k−1, in the frequency subband f, for the partition b.
[0075] Expression of echo-to-signal energy ratio
[0076]
number
[0077] is itself at least the power inter-spectral density between the signal y coming from the microphone and the signal X intended to feed the loudspeaker.
[0078]
number
[0079] can be estimated as a function of
[0080] For example, in matrix representation, where f denotes the row index and b denotes the column index, the expression for the echo-to-signal energy ratio is
[0081]
number
[0082] can be given by:
[0083]
number
[0084] , where β is a positive forgetting factor less than 1, and is written as (k-1) refers to the equation determined in the previous frame (k-1).
[0085] In this formula, the power spectrum density
[0086]
number
[0087] can be given by:
[0088]
number
[0089] , let {α,δ,η,ξ}∈]0,1].
[0090] The power spectral density of the signal intended to feed the loudspeaker X and the power spectral density of the signal y coming from the microphone can be given in matrix notation, where X is a matrix and y is a vector, by:
[0091]
number
[0092] where α and η are forgetting factors greater than 0 and less than 1. Here, |.| 2 The squared norm of a matrix (or vector), denoted as , is defined as the matrix of the norms squared for each element of the matrix.
[0093] In one embodiment that takes advantage of the estimation of the adaptive filter, the adaptive filter may be represented by a contiguous partition. Thus, in such an embodiment, the filter w may be of finite impulse response type and may be N samples long. In particular, the filter may be a filter of L samples each.
[0094]
number
[0095] Partition w b is subdivided into
[0096] In such an embodiment,
[0097]
number
[0098] Like, partition w b which corresponds to the equation in the transformed domain (e.g., in the aforementioned domain of frequency subbands) of
[0099]
number
[0100] Denote the filter in the transformed domain as,F,, where F is the domain transformation matrix.
[0101]
number
[0102] can be estimated.
[0103] In this embodiment, the column index "b" above is now the partition index w b Note that the matrix representation presented above with row index f and column index b may correspond to: However, the matrix representation presented above with row index f and column index b may be applied to situations other than those involving partitions of filters. As a direct illustrative example, the formula given above is still valid in a degraded embodiment, e.g., where b=1 and thus does not involve partitions.
[0104] Furthermore, of the M samples of the signal intended to be fed to the loudspeaker x(t),
[0105]
number
[0106] represents the signal intended to be fed to the loudspeaker for each time frame, denoted as x b =Fx b As,
[0107]
number
[0108] The last B frames x b The matrix corresponding to the transformation of
[0109]
number
[0110] is formed. The time frame of the signal y(t) coming from the microphone
[0111]
number
[0112] Regarding the vector
[0113]
number
[0114] is finally formed.
[0115] This vector y can be constructed as follows:
[0116]
number
[0117] In this form, the acoustic path ΔW for the current frame k (k) The update of
[0118]
number
[0119] where -
[0120]
number
[0121] denotes the Hadamard product, -
[0122]
number
[0123] is expressed as follows: G=FF H and G=I M is a matrix given by either -
[0124]
number
[0125] is the matrix representing the normalization mentioned above, -e (k) is the a priori error estimated from signals x and y for frame k.
[0126] The a priori error may be given by:
[0127]
number
[0128] The adaptive filter is (k) In one embodiment, the update is updated from the current frame k to the next frame k+1 according to an update of , which can be estimated for the current frame k, and the acoustic path update is given by a relationship of the following type: W (k+1) =W (k) +ΔW (k)
[0129] The present specification also relates to a computer program product, which when executed by a processor, comprises instructions for performing the above-described method. In another aspect, there is provided a non-transitory, computer-readable storage medium having such a program stored thereon. The present description also relates to a device for processing a signal y(t) coming from at least one microphone, said device comprising a processor configured to carry out the method defined above.
[0130] Other features, details, and advantages will become apparent upon reading the following detailed description and examining the accompanying drawings. [Brief explanation of the drawings]
[0131] [Figure 1]FIG. 1 illustrates an equipment item in which the subject matter of this document may be implemented, according to one embodiment. [Figure 2] FIG. 1 illustrates a process according to one embodiment for delivering the aforementioned useful signal. [Figure 3] FIG. 1 illustrates a process according to one embodiment for distributing updates to the aforementioned acoustic path estimates. [Figure 4] 1 illustrates a device for implementing the objects of this specification, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0132] Most of the drawings and descriptions below contain elements that are specific in nature, and therefore they not only serve to enable a better understanding of the present disclosure, but also, where applicable, they contribute to its definition.
[0133] In the following, this document proposes an acoustic echo cancellation solution that is robust in double-talk situations, which is generally based on a process involving adaptive filtering, for example an NLMS process, applied successively to each frame of a sequence of frames, a frame being understood here to mean a given number of successive samples of the signal fed to the loudspeaker x(t), this signal being of course assumed to be digital.
[0134] In one embodiment, the filters used for adaptive filtering are partitioned (the length of each partition may or may not correspond to the length of a frame), preferably in the frequency domain (a technique referred to herein as "Partitioned-Block Frequency Domain NLMS" or "PBFD-NLMS"). Techniques of this type are presented, for example, in the reference [@borrallo1992implementation].
[0135] More specifically, the solution here is based on the derivation of the BLUE optimal step size, but estimates the required statistics directly from the reference and microphone signals without adding any auxiliary information. This is similar to the prior art references, specifically [@gil2014frequency], where ΔW is calculated without any error prediction model or a priori error prediction model for the acoustic path. (k) It is possible to calculate
[0136] Such an embodiment guarantees both near-optimal convergence in the sense of speed of convergence, zero bias during convergence, and absence of divergence in double-talk situations, without any auxiliary information other than that directly inferred by the process itself.
[0137] When expressed in the frequency domain, adaptive filtering specifically allows for controlling and normalizing the update of the acoustic path independently of the frequency band involved. Thus, in addition to reduced complexity, the solution benefits from a more uniform convergence over the entire frequency range considered.
[0138] Related to frequency domain filtering, its operation by partitioning also involves filtering the time-frequency representation W of the acoustic path w at each iteration of processing. (k) This allows us to implement different adaptation strategies according to the partition, and it also makes it possible to ensure better convergence in the case of very long filters.
[0139] Such processing makes it possible to derive a step size that optimizes both the behavior in double-talk situations and the acoustic channel tracking.
[0140] Figure 2 shows the different steps of the adaptive filtering solution. At each iteration of the adaptive filtering process, a frame of L new samples of the signals x(t) and y(t) is considered,
[0141]
number
[0142] In step S1, it is determined whether it is necessary to initialize the acoustic paths to be considered (e.g., at the start of a conversation between the speaker and another party), and if so, the initialization of the acoustic paths is performed in step S2. Otherwise, in step S3, the acoustic echo cancellation AEC process is started immediately. In step S4, a time frame of the reference signal x(t) is retrieved, which in the example described is the frequency representation x (k) In step S5, the projection is applied to it in the frequency domain (e.g., in the domain of frequency subbands) to obtain y(t). A similar process is performed for each time frame of the microphone signal y(t) (step S6), to obtain the frequency domain projection y(t) in step S7. (k) Based on the frames x(t) and y(t) (or based on their frequency representation as described herein), the echo cancellation process calculates the prior error e (k) is applied in step S8 to estimate
[0143] An embodiment of adaptive filtering for echo cancellation by generating filters by partitions according to the "partition block" technique is described below.
[0144] has a length of N samples,
[0145]
number
[0146] partitions
[0147]
number
[0148] Consider a target acoustic path modeled by a finite impulse response filter w(t) divided into partitions w(t) as follows: b The matrix corresponding to the frequency transform of
[0149]
number
[0150] can be estimated.
[0151]
number
[0152] , where
[0153]
number
[0154] , M≧L.
[0155] F is the domain transformation matrix, for example, where each element is
[0156]
number
[0157] In practice, the redundancy is achieved by filling with zeros in the time domain.
[0158] Similarly, a time frame contains M samples of the reference signal x(t).
[0159]
number
[0160] Consider the last B frames x b The matrix corresponding to the frequency transform of
[0161]
number
[0162] is formed as follows:
[0163]
number
[0164] , where x b =Fx b Let's say.
[0165] Time frame of microphone signal y(t)
[0166]
number
[0167] By further examining the vector
[0168]
number
[0169] can be shown as follows:
[0170]
number
[0171] To avoid the problems associated with convolution operations performed in the frequency domain, the process is based on overlap-save arithmetic (OLS). Here, the exponent (k) reflects the kth iteration of the process. W (0) , X (0) , y (0)After iterations of the other characteristics, the process
[0172]
number
[0173] We can continue by calculating
[0174]
number
[0175] however,
[0176]
number
[0177] denotes the Hadamard product here, and (.) * denotes the conjugate of a matrix or vector.
[0178] As mentioned above, the redundancy of the DFT is found in the a priori error formula and is realized by zero-padding, which advantageously makes it possible to avoid artifacts due to circular convolution.
[0179] The method then proceeds to step S9, where the acoustic path
[0180]
number
[0181] Continue computing updates to
[0182] That is,
[0183]
number
[0184] , where
[0185]
number
[0186] Let's say.
[0187] In an embodiment where the update is optimal (due to zero-padding), G=FF H (called a "constrained" update).
[0188] In embodiments where the update is suboptimal, G=I instead M (called "unconstrained" update). Such an implementation has the advantage of consuming fewer resources.
[0189] Then, step S10 is performed to update the acoustic path by (k+1) The aim is to calculate W (k+1) =W (k) +ΔW (k)
[0190] and in step S11, the useful signal
[0191]
number
[0192] is W (k) By x (k) is obtained after convolution with and is returned to the time domain.
[0193] Figure 3 shows the acoustic path W (k) This section details the steps for updating Λ, specifically the calculation of the optimal regularization term Λ, which allows for inherent robustness in double-talk situations.
[0194] To ensure that the echo cancellation solution is robust in the situations described above, a spectral normalization term Λ that satisfies the BLUE criterion is selected. (k) is chosen. This can be achieved by knowing the power spectral density (PSD) of the microphone signal x(t) and of the local signal s(t).
[0195] In reference [@van2007double], BLUE is obtained at the cost of strong assumptions about the local signal, which must be accepted using an autoregressive model considered to be speech and an error prediction method. On the other hand, [@trump1998frequency] obtains BLUE by using the error signal e (k) We achieve BLUE by estimating the PSD of the local signal a posteriori only, doing so at the cost of reduced stability, and operating with a strong constraint on the local signal (stationary colored noise).
[0196] The proposed solution, described below, overcomes the constraints used above by robust estimation of the DSP over time.
[0197] Assume the following: -
[0198]
number
[0199] (resp.
[0200]
number
[0201] ) the power spectral density (PSD) estimate of X at each frequency and (resp. at each frequency) at each partition of y, - Between the spectrum of the microphone signal and the reference signal
[0202]
number
[0203] and its power spectrum density
[0204]
number
[0205] , The power spectral density of the local signal s at each frequency and each partition is
[0206]
number
[0207] It is shown as follows.
[0208] Finally, for each frequency band and each partition, a matrix representing the ratio of the energy of the echo and the local signal is defined as
[0209]
number
[0210] (denoted ESR for "Echo-to-Signal Ratio").
[0211]
number
[0212] After initializing the and other indices, the process performs the following operations:
[0213] - Power Spectral Density (PSD) estimation:
[0214]
number
[0215] ,however
[0216]
number
[0217]
number
[0218] ,however
[0219]
number
[0220]
number
[0221] Here, let {α,δ,η,ξ}∈]0,1].
[0222] - Instantaneous echo-to-signal ratio (ESR) estimation:
[0223]
number
[0224] , Here, let β∈]0,1].
[0225] - normalization Λ (k) Estimate of:
[0226]
number
[0227] Then, by applying the method within the meaning of this description, the normalization parameters to meet the BLUE criteria are expressed by:
[0228]
number
[0229] , where μ∈[0,2[.
[0230] This last expression for the normalization parameter finally includes the term ∑ k = ...
[0231]
number
[0232] Includes:
[0233] Note that in some cases, to implement the traditional technique of PBFD-NLMS, by instead applying state-of-the-art teachings, e.g., as described in [@borrallo1992implementation], the regularization parameters of the filters can be expressed as:
[0234]
number
[0235] , where μ∈[0,2[ and do not include any measurements in the estimation of the echo-to-signal ratio.
[0236] 3, a first step S20 begins with a test to determine whether the power spectral density estimation is initialized. If so, in step S21, the respective spectral densities of the signal from microphone y and the reference signal x are initialized. If not, a procedure for estimating the spectral normalization factor Λ is immediately initiated in step S22. In step S23, the current frequency frame of the microphone signal is retrieved, and in step S24, the current frequency frame of the reference signal is retrieved in order to estimate the aforementioned inter-spectral density in step S25. Then, in step S26, the power spectral density of the microphone signal is estimated, and in step S27, the power spectral density of the reference signal is estimated in order to deduce therefrom an estimate of the instantaneous echo-to-signal ratio (ESR) in step S28, as explained above. The spectral normalization factor Λ (k) An estimate of is deduced from this in step S29, from which an acoustic path update can be determined in step S30.
[0237]
number
[0238] Achieving the BLUE criterion using adaptive filtering performed in the time domain involves the following solution:
[0239]
number
[0240] is equivalent to finding
[0241]
number
[0242] ,however
[0243]
number
[0244] is the microphone signal vector,
[0245]
number
[0246] is the matrix of loudspeaker signals,
[0247]
number
[0248] is the autocorrelation matrix of the signal s.
[0249] Current echo cancellation methods are based on frequency domain adaptive filtering. The method presented in [@Trump1998frequency] filters the frequency domain acoustic channel as follows:
[0250]
number
[0251] We propose to achieve a regularized version of the BLUE criterion by searching for a solution.
[0252]
number
[0253] , where Γ s (resp.Γ x ) is a diagonal matrix of the power spectral density of the signal s(resp. x).
[0254] However, the local signal s is unknown, and an estimator that satisfies the BLUE criterion is therefore extremely difficult to obtain in practice without other information or a model for s.
[0255] The local signal s and the reference signal x are decorrelated, i.e., E{Xs T If}=0 (let E{·} denote the expectation operator), the solution given in [@Borrallo1992implementation] simply
[0256]
number
[0257] can produce an unbiased estimate of (and therefore satisfy the BLUE criterion for) . In practice, this condition may only be met if the signal s is white noise. In all other situations, the unbiased estimate
[0258]
number
[0259] To reach this, the normalization factor Λ (k) In the denominator, the variance part of the local signal s, that is, E{ss H} or its power spectral density (PSD) (parameter
[0260]
number
[0261] (which takes the form of
[0262] Expected value E{ss H} are unknown, the solution proposed here is to use the power spectral density, and in particular the denominator as done above, i.e., Γ s, overcomes the constraints used above by a time-dependent robust estimation of the power spectral density of the local signal s.
[0263] The process described above can be used in particular in situations where it is necessary to capture sound and play it back at the same time: the most common use cases are hands-free telephony (a person speaking at a distance hears a delayed version of their own voice mixed with the voice of the other person, i.e. an echo), dialogue with voice assistants (replies from the dialogue system and / or music played by the voice assistant mix with the commands issued by the user and interfere with voice recognition), intercoms, video conferencing systems, etc.
[0264] A device for implementing the above method is represented in FIG. 4 and may be indicated by the two modules (adaptive filtering and subtraction applied to the signal y(t) captured by the microphone) on the left side of FIG. 1. Referring to FIG. 4, this device may generally comprise a first input interface IN1 for receiving the signal y(t) collected from the microphone MIC, and a second input interface IN2, which in the illustrated example is for receiving a signal to be reproduced on the loudspeaker HP (e.g., a telecommunications signal such as a voice or music signal). The device comprises a processor PROC that can cooperate with a memory MEM to process this audio signal and deliver a signal x(t) intended to be supplied to the loudspeaker HP via a first output interface OUT1 provided in the device. In particular, the memory MEM stores instruction data of a computer program according to at least one aspect of this description, which instruction data is readable by the processor PROC to perform the above-described processing and, in particular, apply it to the signal from the microphone y(t) in order to deliver a useful signal s(t) via a second output interface OUT2 provided in the device in one exemplary embodiment.
[0265] Naturally, this is an exemplary embodiment, and here in general the useful signal s(t) may be transmitted via the output interface OUT2, for example to a distant party. In this case, the interface OUT2 may be connected, for example, to a communications antenna or a router of a telecommunications network NET. The same applies to the input interface IN2, which receives a signal "from the outside" to be played through a loudspeaker.
[0266] For example, in a device such as a voice assistant, it is common to encounter double-talk situations in which the user is speaking a voice command at the same time that the assistant is, for example, responding to a previous command. In this case, at least some of the voice assistant's responses can be generated locally from the contents of the memory MEM, for example, without the need to utilize a remote server and telecommunications network. Additionally, the useful signal s(t) can simply be interpreted locally by the processor PROC to respond to a voice command from the user. The interfaces IN2 and OUT2 may therefore not be necessary.
[0267] A typical use case for so-called "double talk" processing with a voice assistant consists, for example, of a user listening to music through the voice assistant's loudspeaker while speaking a command to wake up the assistant (WakeUpWord). In this case, to be able to correctly detect the actual spoken command signal s(t), the signal of the playing music x(t) must be extracted from a sound signal y(t) captured in the exact environment where the assistant is located (with its reverberations). * It is wise to eliminate w(t).
[0268] Naturally, the invention is not limited to the embodiments presented above, but can extend to other variants.
[0269] For example, practical embodiments have been described above in which signals are processed in the domain of frequency subbands. However, the present invention also provides a method for processing the power spectral density Γ of a local signal s in the frequency domain. s The expected value E{ss T} may alternatively be implemented in the time domain by utilizing parameters such as
[0270] As a result, the foregoing estimation of the power spectral density can be performed effectively in one possible, but not required, implementation.
[0271] The same is true for adaptive filter partitions, which can also operate in the subband domain, although this is not required.
[0272] 4 shows a compact item of equipment comprising an echo cancellation device (which may therefore be represented by a processor PROC, a memory MEM, at least one input interface and at least one output interface), as well as a microphone MIC and a loudspeaker HP. In a variant, the device, the microphone(s) on the one hand, and the loudspeaker(s) on the other hand, may be located in different places and connected, for example, by a telecommunications network, or a local area network (operated by a home gateway), or by other means.
[0273] Attachment: References [@Borrallo1992implementation]: Borrallo, JP and Otero, MG (1992). On the implementation of a partitioned block frequency domain adaptive filter (PBFDAF) for long acoustic echo cancellation. Signal Processing, 27(3), 301~315
[0274] [@Trump1998frequency]: Trump, T. (1998, May). A frequency domain adaptive algorithm for colored measurement noise environment. In Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP'98 (Cat. No. 98CH36181) (Vol. 3, pp. 1705 - 1708). IEEE
[0275] [@Jung2005new]: Jung, H. K., Kim, N. S., and Kim, T. (2005). A new double - talk detector using echo path estimation. Speech communication, 45(1), 41 - 48
[0276] [@van2007double]: Van Waterschoot, T., Rombouts, G., Verhoeve, P., and Moonen, M. (2007). Double - talk - robust prediction error identification algorithms for acoustic echo cancellation. IEEE Transactions on Signal Processing, 55(3), 846 - 858
[0277] [@gil2014frequency]: Gil-Cacho, JM, Van Waterschoot, T., Moonen, M., and Jensen, SH (2014). A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 22(12), 2074~2086 [Explanation of symbols]
[0278] f frequency, frequency subband HP loudspeakers IN1 First input interface IN2 Second input interface k instants, frames MEM memory OUT1 First output interface OUT2 Second output interface PROC Processor s Local signal s(f) useful signal s(t) signal, local signal w Acoustic channel, filter w(t) sound path x(t) Microphone signal, reference signal y(t) microphone signal z(t) echo signal
Claims
1. 1. A method for processing a signal y(t) coming from at least one microphone (MIC) of an item of equipment, said item of equipment further comprising at least one loudspeaker (HP) intended to be supplied with a signal x(t), Processing the signal y(t) from the microphone (MIC) A filter is applied to the signal x(t) fed to the loudspeaker (HP) to limit echo effects. [Equation 1] The echo signal given by applying [Equation 2] Determination of the useful signal s(t) by subtracting the estimate of from the signal y(t) coming from the microphone (MIC) [Equation 3] Including, The filter [Equation 4] is adaptable with a variable step size to take into account changes over time in the acoustic path w(t) from the loudspeaker to the microphone; the signal x(t) supplied to the loudspeaker is obtained in the form of a time-dependent succession of frames of signal samples; The adaptive filter [Equation 5] is the update ΔW of the acoustic path w(t) for each frame k of samples. (k) is generated at this frame k by applying a normalization Λ that satisfies the chosen criterion of minimum variance according to the normalization Λ is a function of a parameter representing the statistical expectation of the useful signal s(t), the adaptive filter is generated in the domain of a frequency subband f; The normalized Λ is power spectral density of said useful signal s [Equation 6] and, power spectral density of the signal x fed to the loudspeaker [Equation 7] and is defined as a function of the power spectral density of the useful signal s [Equation 8] is the power spectral density of the signal y captured by the microphone [Equation 9] , and an expression for the echo-to-signal energy ratio [Equation 10] ,method, which is estimated as a function of .
2. The method of claim 1 , wherein the chosen criterion is a BLUE type of best linear unbiased estimator.
3. In the matrix representation, where f denotes the row index and b denotes the column index, the normalized Λ (k) (f,b) is [0011] is given by 3. The method of claim 1 or 2, wherein μ∈[0,2[, where γ is a chosen positive coefficient.
4. In matrix representation, where f denotes the row index and b denotes the column index, the power spectral density of the useful signal s [0012] but, [0013] is given by where A is the chosen positive limit, [0014] 4. The method according to claim 1, wherein s is the power spectral density of the useful signal s evaluated for the previous frame k-1, in frequency subband f, for partition b.
5. the representation of the echo-to-signal energy ratio [Equation 15] is at least the power spectral density between the signal y coming from the microphone and the signal X intended to feed the loudspeaker [0016] 5. The method of claim 1, wherein the value of the sigma-based ...
6. the expression of the echo-to-signal energy ratio in a matrix representation where f denotes a row index and b denotes a column index. [Equation 17] but, [Equation 18] is given by where β is a positive forgetting factor smaller than 1, and is expressed as (k-1) The method of claim 5 , wherein refers to the equation determined in the previous frame (k−1).
7. the power spectrum density [Equation 19] but, [Equation 20] is given by The method of claim 6, wherein {α,δ,η,ξ}∈]0,1].
8. a signal intended to be fed to said loudspeaker, represented by a matrix X; and The signal coming from said microphone is represented by the vector y The power spectral densities of [Equation 21] , and [Equation 22] is given by 8. The method according to claim 6, wherein α and η are forgetting factors greater than 0 and less than 1.
9. The adaptive filters are N samples long and each have L samples. [Equation 23] Partition w b 9. The method of claim 1, wherein the finite impulse response filter w is subdivided into [Request Item 10] [Number 24] As in, the partition w b corresponds to the transformed domain expression of [Equation 25] where F is the domain transformation matrix, and the matrix [Equation 26] Estimate of M samples of the signal x(t) intended to be fed to the loudspeaker, [0000] For each time frame, denoted as x b =Fx b Then, [0000] The last B frames x b Corresponding to the transformation of the matrix [0000] is formed, The time frame of the signal y(t) coming from the microphone [Equation 30] Regarding the vector [Equation 31] The method of claim 9, wherein
11. The vector y is [Equation 32] The method of claim 10, wherein
12. The update ΔW(k) of the acoustic path w(t) for the current frame k is [Equation 33] where [Equation 34] denotes the Hadamard product, [Equation 35] is expressed as follows: G=FF H and G=I M is a matrix given by either [Equation 36] is the matrix representing the normalization mentioned above, e (k) 12. The method of claim 10 or 11, wherein ∑ k = ...
13. The prior error is [Equation 37] The method of claim 12, wherein the formula is given by:
14. The adaptive filter is configured to satisfy a relationship of the type W (k+1) =W (k) +ΔW (k) 14. The method according to claim 1, wherein the acoustic path W(k) is updated from the current frame k to the next frame k+1 according to an estimated update ΔW(k) of the acoustic path W(k) for the current frame k according to:
15. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any one of claims 1 to 14.
16. 15. A device for processing a signal y(t) coming from at least one microphone (MIC), said device comprising a processor (PROC) configured to carry out the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Speech-processing device, speech-processing method, and information-processing device
WO2019044176A1