Attention mechanism for estimating a delay and the alignment of two signals
The method employs a trained filtering module to generate weighted attention scores for accurate delay estimation and alignment between two signals, addressing the limitations of current echo cancellation techniques in variable environments.
Patent Information
- Application Number
- PCT/EP2024/086947
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
Current echo cancellation techniques struggle to accurately estimate delays between signals in variable environments, especially when echo is absent, exhibits complex temporal variations, or is interfered with.
A method and system for temporal alignment between two signals using a trained filtering module that generates weighted attention scores, allowing for accurate delay estimation and alignment, even in complex or noisy environments.
The proposed method provides a robust and reliable estimate of the delay between signals, improving echo cancellation and signal alignment in various contexts, including those with intermittent or absent interference.
Smart Images

Figure EP2024086947_26062025_PF_FP_ABST
Abstract
Description
Attention mechanism for delay estimation and alignment of two signals
[0001] This disclosure is in the field of information and communications technologies, and more specifically in the field of signal processing.
[0002] It relates to a method and a system for estimating a delay between an expression of a first signal and an expression of a second signal, a computer program for implementing such a method and a non-transitory recording medium readable by a computer on which such a program is recorded.
[0003] Accurately estimating the delay between two potentially coupled signals is a common problem in several technical fields. Current solutions are often designed for specific use cases.
[0004] For example, current voice communication systems, especially those used in hands-free environments, frequently face the problem of echo. This echo is mainly due to impedance differences in wired connections and acoustic or structure-borne coupling between loudspeakers and microphones. Traditional echo cancellation techniques involve the use of adaptive but also non-linear filters operating in the frequency domain.
[0005] Although current methods, including those using neural networks for time-frequency mask estimation, have improved the performance of echo cancellation, they remain insufficient in certain scenarios. In particular, they struggle to effectively handle situations where the echo is absent, exhibits complex temporal variations, or when interference is added to the echo.
[0006] There is a need for a robust method capable of accurately estimating a delay between two signals in a variety of contexts, particularly in variable environments where interference between these two signals may be intermittent or absent. Summary
[0007] The invention aims to improve the situation.
[0008] According to one aspect, there is provided a method for temporal alignment between characteristics of a first signal and characteristics of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the method comprising, for a current frame of the second signal:obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, andobtaining, by a trained filtering module, a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally neighboring the current frame,weighting the attention scores with the weights obtained to produce a set of weighted attention scores;a linear combination of the features of the first signal with the weighted attention scores to obtain an estimate of the realigned features.;
[0009] According to one aspect, there is provided a system for aligning features of a first signal with features of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the system comprising the following modules applied to a current frame of the second signal: a module for obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, a trained filtering module capable of - obtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally adjacent to the current frame, and - weighting attention scores with the weights obtained to produce a set of weighted attention scores;a combination module capable of linearly combining features of the first signal with the weighted attention scores to obtain an estimate of the realigned features.;
[0010] A technical advantage of the proposed method and system lies in their ability to provide an accurate and reliable estimate of the overall delay between features of the first signal and features of the second signal and thus of the realigned features. This is particularly beneficial in contexts where the overall delay between expressions of the first signal and the second signal remains relatively constant from one frame to another. In such contexts, constancy or limited variation of the sets of weighted attention scores from frame to frame indicates stability of the estimated overall delay.
[0011] In one example, the characteristics of the two signals are normalized so that their norm is 1. This makes the system independent of the difference in scale factor between the two signals.
[0012] In one example, the resulting set of weighted attention scores is normalized to correspond to a set of probabilities.
[0013] This normalization makes it possible to provide consistent information over time, from one frame to another, for processing modules located downstream. According to one aspect, the method further comprises detecting characteristics of the first signal in the second signal by determining a probability of presence of a characteristic of the first signal in the second signal on the basis of processing the sets of weighted attention scores comprising an implementation of one-dimensional convolutional layers acting on a time axis of delays, each convolutional layer being configured to perform a sub-sampling of the time axis of delays and an implementation of a pooling layer to deduce the probability of presence therefrom.
[0014] Detection of the expression of a first signal within a second signal is improved by processing weighted attention score sets, which significantly reduces false detections, resulting in more robust alignment.
[0015] This implementation offers the advantage of a fine and detailed analysis of the signals, allowing precise localization of the characteristics of the first signal within the second signal and reliable detection even in the presence of complex or noisy signals.
[0016] In one example, the combination of the weighted attention score sets includes, after the implementation of the convolutional layers, an implementation of a pooling layer to infer the probability of presence. This pooling layer is configured to make the probability of presence of the expression of the first signal within the second independent of the number of frames forming the given temporal context. This makes it possible to change the size of the temporal context considered without having to retrain the neural network.
[0017] In one embodiment, the method further comprises a step of smoothing the probabilities obtained by implementing a recursive neuron. This makes it possible to avoid overly rapid variations which are not very representative of real conditions.
[0018] Also provided, in one aspect, is a method of canceling an echo represented by characteristics of a first signal within a second signal, the method comprising aligning the characteristics of the first signal with the characteristics of the second signal according to the proposed alignment method.
[0019] Effective echo cancellation requires an accurate estimation of the delay and alignment between the expression of the first signal to be canceled and the expression of a second signal containing a useful signal to be preserved. This accurate estimation is provided by the use of weighted attention scores. Furthermore, the implementation of the proposed detection allows, thanks to the processing of sets of weighted attention scores, to efficiently detect the presence or absence of echo and, consequently, to preserve the quality of a useful signal by avoiding processing a non-existent echo.
[0020] Also provided, according to one aspect, is a method of determining an origin of a first signal and a second signal, the first signal and the second signal being sensed distinctly, based on a delay between characteristics of the first signal and characteristics of the second signal, the method comprising an alignment between the characteristics of the first signal and the characteristics of the second signal according to the proposed method.
[0021] This method offers the advantage of accurately tracing the source of signals in the event of multiple transmission paths or reflections, which is essential in applications such as medical diagnosis by ultrasound, geolocation and structural diagnosis by seismography techniques.
[0022] The alignment process can also be called a process of synchronizing a first signal and a second signal.
[0023] This process offers the advantage of precise coordination between signals, which is essential in multimedia applications such as lip synchronization in video conferencing or synchronization of audio tracks in film post-production, where even a minimal time shift can be noticeable and annoying.
[0024] Also provided, in one aspect, is a computer program comprising instructions which, when the program is implemented by a processor, result in implementing one of the methods discussed.
[0025] Also provided, in one aspect, is a non-transitory computer-readable recording medium having recorded thereon a program for implementing one of the methods discussed when that program is executed by a processor.
[0026] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which: Fig. 1
[0027] represents a state-of-the-art acoustic echo cancellation algorithm. Fig. 2
[0028] represents an attention model applied to acoustic echo cancellation according to the state of the art. Fig. 3
[0029] represents an attention model applied to acoustic echo cancellation according to an exemplary embodiment. Fig. 4
[0030] represents an attention score filtering module and an echo presence detector suitable for the attention model of. Fig. 5
[0031] represents an example of the result of a time alignment of two signals in the absence of echo, the time alignment using a delay determined by application of the model of. Fig. 6
[0032] represents an example of the result of a time alignment of two signals in the absence of echo, the time alignment using a delay determined by application of the model of. Fig. 7
[0033] represents an example of the result of a time alignment of two signals in the presence of interference and echo, the time alignment using a delay determined by application of the model of. Fig. 8
[0034] represents an example of the result of a time alignment of two signals in the presence of interference and echo, the time alignment using a delay determined by application of the model of. Fig. 9
[0035] represents an example of a device for estimating the delay between a first signal and a second signal.
[0036] In the following description, like reference numerals designate identical elements or elements having similar functions.
[0037] La description qui va suivre mentionne les références suivantes.
[0038] [Indenbom2022] Evgenii Indenbom, Nicolae-Cătălin Ristea, Ando Saabas, Tanel Pärnamaa, and Jegor Gužvin, “Deep model with built-in self-attention alignment for acoustic echo cancellation,” Aug. 2022, arXiv:2208.11308 [cs, eess].
[0039] [LeRoux2019] J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr - half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.
[0040] [Fu2019] Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, Shou-De Lin, “MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement”, International Conference on Machine Learning, 2019.
[0041] The following description refers to known technologies in the field of signal processing, in particular known artificial intelligence technologies.
[0042] A presentation of a case study in which such technologies are used for echo cancellation is now provided to facilitate the understanding of the proposed technique.
[0043] During voice communication, especially hands-free, an echo phenomenon can appear. Several causes are at the origin of this echo. Wired hybrid connections cause impedance differences, sources of audible reflections like an echo with a certain delay. On the terminal side, couplings exist between the speaker and the microphone. These couplings can be acoustic in nature, especially in the case of hands-free communications: reverberation or room effect, due to the walls, create multiple reflections that will be picked up by the microphone. There are also solid couplings within the terminal due to the vibration of certain parts, vibrations picked up by the microphone.Due to these couplings, the voice of the remote interlocutor broadcast by the loudspeaker is picked up by the microphone with a certain delay and sent back to the remote party, creating an echo phenomenon which degrades the communication and impacts the intelligibility of the local interlocutor. Also, telephone systems integrate echo cancellers at different points in the chain in order to eliminate this artifact.
[0044] illustrates an echo canceller (102) configured to cancel, or at least attenuate, an echo representing an expression of a signal emitted by a loudspeaker (104) within a signal picked up by a microphone (106).
[0045] The following notations are introduced: d(t) denotes the signal emitted by the loudspeaker, ϕ(.) denotes the coupling between the loudspeaker and the microphone, e(t) denotes the echo picked up by the microphone, s(t) denotes the useful signal, and x(t) denotes the signal picked up by the microphone.
[0046] x(t) is given by the following additive relation, denoted [1].x(t)=s(t)+e(t)=s(t)+ϕ(d(t))
[0047] The coupling ϕ(.) between the loudspeaker and the microphone is characterized by: linear transformations representative of the frequency response of the transducers (loudspeaker / microphone) and the reverberation of the room, non-linear transformations, such as saturations that can appear at the level of the transducers, and an overall delay, noted τ, which represents the sum of the delay of the acquisition-restitution chain of the terminal (jitter buffers, digital-analog conversions) and the acoustic delay linked to the propagation of sound in the air.
[0048] Echo cancellation techniques typically use a serial combination of an adaptive filter seeking to estimate the linear transformation and a non-linear filter seeking to suppress the residual echo at the output of the adaptive filter. For computational and performance reasons, processing is generally carried out in the frequency domain: this allows, on the one hand, to design more robust and faster adaptive filters, and on the other hand, to develop non-linear filters in the form of time-frequency masks that guarantee a better compromise between the suppression of the residual echo and voice quality.
[0049] Thus, each signal noted z(t), where z(t) designates one of the signals introduced previously, notably d(t) and x(t), is decomposed into frames – adjacent or not – before applying a short-term Fourier transform to produce a frequency vector noted Z(n) ∈ ℂ F, that is to say D(n) or X(n) respectively, where n is the frame index and F is the number of frequencies considered. First, an apodization window noted w∈ ℝ N , where N is a natural integer, is applied to each signal frame to control the effects of windowing. The apodization window can for example be a Hann, Hamming, or Blackman window.
[0050] The application of the apodization window can be formalized by the following equation:Z(n) = ℱ(z(n)⊙w)wherez(n) is a vector of size N representing the frame of index n of the signal z(t),ℱ(.) denotes a short-term Fourier transform, and⊙ denotes a Hadamard product, or term-by-term product, of two vectors or matrices.
[0051] The overlap rate of two successive frames, noted (NP) / N, is fixed by the number P. P is a natural integer between 1 and N. When P=N, the overlap rate is equal to zero, indicating that the frames are adjacent.
[0052] The terms of the vector z(n) are noted z(n) = [z(n*P) z(n*(P-1)) ... z(n*(P-N+1))].
[0053] Under this formalism, the non-linear module generally has the functions of: calculating a time-frequency mask m(n) ∈ ℝ F apply this mask to the output signal of the adaptive filter to estimate the useful signal in the time-frequency domain, noted Ŝ(n) and reconstruct the time version of the estimated useful signal, noted ŝ(t), by inverse Fourier transform and by application of known methods of the “overlap-and-add” or “overlap-and-save” type.
[0054] In recent years, the use of deep neural networks to estimate the maskm(n) has shown very interesting performances [Indenbom2022], in any case much superior to signal processing techniques based on the estimation of Wiener filters from signal models (probabilistic models) and the coupling function ϕ(.). The very latest approaches are sufficiently robust to no longer use an adaptive filter.
[0055] These approaches involve applying the mask directly to the microphone signal. This application can be formalized by the following equation, denoted [2].Ŝ(n)=m(n)⊙X(n)
[0056] In the latter case, the transform can be learned by the neural network. The architecture of the neural network then includes an encoding step, denoted Γ(.) implemented by a series of layers of neurons whose role is to extract characteristics of the signal considered, namely the signal emitted by the loudspeaker and the signal captured by the microphone, respectively.
[0057] The encoding performed by the neural network can be formalized by the following equation, considering Q as being the set of indices of past or future frames to be provided as input to the encoder: Z(n) = Γ({z(nq)} q ∈Q )
[0058] In the remainder of this description, the use of the notation Z(n) refers to a frequency vector resulting, indifferently, from a short-term Fourier transform or from an encoding step carried out by the neural network.
[0059] To estimate the maskm(n) of the frame of index n, the neural network is fed with the features ofX(n) and with the features ofD(ni).
[0060] The index i here delimits the relevant temporal context, noted D, for the frame of index n.
[0061] This temporal context is defined by a set of frames encompassing both past and future frames relative to the frame of index n, expressed as {,-n_future ≤ i < n_past}. In other words, for each frame of index n considered, the neural network takes into account a set of neighboring frames, extending from n_past frames in the past to n_future frames in the future.
[0062] The mask estimation is formalized by the following equation noted [3].m(n) = Ψ(X(n),D(n))
[0063] Ψ(.) is a mask generation function. Its implementation can take the form of an algorithm or all or part of a neural network.
[0064] Conventionally, due to the overall delay τ, echo cancellation systems integrate a time realignment module which consists of aligning the loudspeaker signal with the microphone signal by compensating for this delay, which allows better conditioning of the different estimators (adaptive filter and / or echo suppressor). Thus, the input of the echo cancellation system becomes d(t-τ), where τ is an estimate of the delay τ, and all the loudspeaker characteristics are calculated from this re-phased signal. In the case of neural networks, rather than explicitly estimating the delay τ, the network integrates in its architecture a realignment module which exploits a set, denoted {D(ni)} i ∈ D , characteristics of D(ni). The realignment by the realignment module is formalized by the following equation noted [4].m(n) = Ψ(X(n),{D(ni)} i ∈ D )
[0065] The temporal contextDest must be determined according to different constraints. The maximum past n_past to be taken into account depends on the overall delay τ and the duration of the room reverberation. In practice, n_past can take a value corresponding to several hundred milliseconds. As for the value of n_future which fixes the quantity of future that we allow ourselves to take, it depends on the causality constraint that we impose on ourselves: if we want processing without delay, as for example for telephony applications, we can set n_future = 0; in other situations where the strict causality constraint is relaxed, we can take one or more frames from the future to better estimate the maskm(n) of the current frame of index n.
[0066] In [Indenbom2022], the authors use an attention mechanism to achieve this realignment. The principle of the attention mechanism applied to echo cancellation consists of linearly combining the set {D(ni)} i ∈ D characteristics of D(ni) with weights noted a(n,i) so as to obtain an estimator of the realigned characteristics, noted D r (n), according to the following equation noted [5].D r (n)=∑ i ∈ D a(n,i).D(ni)
[0067] The weights a(n,i), called attention scores and calculated at each frame of index n, are obtained using a standard attention model which can be formalized according to the following equation, noted [6].a(n) = softmax((1 / √b)X p (n) T [D p (n+n_future), … ,D p (n-n_past)])
[0068] Where a(n) is the vector containing the a(n,i),X p (n) =W q T .X(n) and D p (n) =Wq T .D(n) whereW q ∈ ℝ cxb andW k ∈ ℝ vxb are matrices called "query" and "key" with c and v the dimensions of X(n) and D(n), respectively, and b the size of the space into which the expressions of the signals are projected.
[0069] The above attention mechanism that can be implemented within an echo canceller is illustrated in.
[0070] A first matrix multiplication operator (202) performs the product of the vector X(n) T and the “query” matrix W q .
[0071] For each vector D(ni), a second corresponding matrix multiplication operator (204a, 204b, 204c) performs the product of the “key” matrix W k T and the aforementioned vector D(ni).
[0072] For each vector D(ni), a third corresponding matrix multiplication operator (206a, 206b, 206c) performs the product of the output vector W k T D(ni) of the second corresponding operator and the output vector X(n) T W q of the first operator.
[0073] An attention module (208) based on a neural network and comprising a softmax activation function receives as input the outputs of the third operators and calculates as output a set of attention scores a(n,i), i.e. an attention score for each vector D(ni).
[0074] For each vector D(ni), a corresponding multiplication operator (210a, 210b, 210c) performs the product of the corresponding attention score a(n,i) and the aforementioned vector D(ni).
[0075] An addition operator (212) sums the outputs of the multiplication operators (210a, 210b, 210c) and thus obtains the estimator of the realigned characteristics according to equation [5].
[0076] In the context of time alignment, we expect a(n,i) to be maximum, i.e. close to 1, around the index noted i τ corresponding to the overall delay τ, and takes a low value, i.e. close to 0, elsewhere.
[0077] The index i τ is calculated as follows, as being the integer part of the quotient of τ by P:i τ =|τ / P|
[0078] In light of the case study as provided above, the proposed technique is now presented in the context of an application to echo cancellation.
[0079] In this description the following terms used are defined as follows:The term "expression of a signal" refers to a representation or manifestation of a signal, which may be a temporal, frequency or any other transformed form of this signal (for example, a spectrogram or a representation in a latent space). We may also speak of characteristics of the signal.
[0080] The term “attention score” refers, in the context of the application, to a value calculated by an attention module (a process known in the field of deep learning) and representing the probability of matching between frames of two signals. The term “unweighted attention score” means the initial or raw score as directly determined by the attention module and before the application of any weighting or adjustment based on additional contexts or weights.
[0081] The term "specific delay" refers to a precise delay calculated for each current frame of the second signal relative to the frames of the first signal. The term "estimated overall delay" refers to the aggregation of these specific delays to give an overall estimate of the delay between the two signals.
[0082] The term "time-neighboring frames" of the current frame means the frames that directly surround the current frame, including, for example, between one and ten frames just before and between zero and ten frames just after it in a time sequence.
[0083] The term "delay axis" represents the temporal dimension used to align the frames of the two signals according to the calculated delays. It is a conventional reference in signal processing systems for timing analysis.
[0084] The proposed technique is illustrated in and is to be compared with.
[0085] The results obtained by applying the proposed technique to test scenarios are illustrated in and can be compared to the results obtained by applying an attention mechanism according to, illustrated in and.
[0086] , [Fig.. 6],and each comprise four representations, namely:a representation (502, 602, 702, 802) of the characteristics of the signal emitted by the loudspeaker as a function of time,a representation (504, 604, 704, 804) of the characteristics of the signal picked up by the microphone as a function of time,a representation (506, 606, 706, 806) of the characteristics of the signal emitted by the loudspeaker synchronized by the system with the signal picked up by the microphone as a function of time, anda representation (802, 806) of the attention scores as a function of time enetand a representation (804, 808) of the weighted attention scores as a function of time enet.
[0087] The proposed technique provides in this context a time alignment solution robust to interference, and to the presence or absence of echo in the signal captured by the microphone. The proposed technique makes it possible in particular to obtain, contrary to the behavior observed by a classic attention module, scores which reflect desired characteristics for a delay estimator between the signal captured by a microphone and the signal emitted by a loudspeaker and an expression of which is likely to be present in the signal captured by the microphone.
[0088] In particular, according to the proposed technique, each frame of index n is associated with a set of scores which can be represented by a vector noted â(n) = [â(n,n_future ), … , â(n,n_past)].
[0089] This set of scores has values concentrated around a particular delay corresponding to the overall delay between the speaker and the microphone.
[0090] In the continuous presence of echo, the set of scores associated with each frame is stable over time, that is to say from frame to frame, including in the presence of a useful signal s(t) captured by the microphone.
[0091] Furthermore, unlike a standard attention mechanism, the proposed technique allows to efficiently model echo-free scenarios. In particular, the vector â(n) can be designed to be representative not only of a delay estimate, but also of a presence or absence of coupling between the signal emitted by the loudspeaker and the signal picked up by the microphone. For example, the vector â(n) can be designed to have a variable norm depending on the presence or absence of such coupling. For example, the vector â(n) can be designed to have a norm close to zero when the coupling between the loudspeaker signal and the microphone signal is zero (ϕ(.)=0).
[0092] Indeed, the standard attention model is originally designed for application to natural language processing and has properties that are not ideally suited to application to the temporal realignment of audio signals in the context of echo cancellation.
[0093] The standard attention model constrains the attention scores to sum to one, through the use of the softmax function in equation [6]. This constraint is problematic because, in the context of echo cancellation, the coupling between the signal emitted by the loudspeaker and the signal picked up by the microphone can be zero, i.e. ϕ(.)=0. In the absence of echo, it is not desirable for the model to identify a correlation between the microphone and the loudspeaker. Indeed, the presence of non-zero values a(n,i) amounts to feeding the echo canceller with features from the signal coming from the loudspeaker, which indicates to the echo canceller the presence of echo and leads in practice to attenuate, and therefore potentially degrade, a local speech signal s(t)=x(t) when no echo is present.
[0094] This case is illustrated where the loudspeaker emits a signal that is not present at the microphone because the coupling is zero.
[0095] We observe that the standard attention mechanism, represented in (508), still identifies links between X(n) and characteristics of the set {D(ni)} i ∈ D characteristics of D(ni). In this case, the identified links do not reflect a delay estimate but simply indicate to the echo canceller that it is still necessary to take into account certain characteristics D(ni) to calculate the mask m(n).
[0096] In an application of the standard attention model to natural language processing, the attention score can vary greatly from one moment to the next. However, in an application to echo cancellation, the delay rarely varies and it is desirable that a(n,i).D(ni) varies little from frame to frame. However, we observe in practice that with standard attention models, if the interference, expressed in the form of a frequency vector S(n) (like X(n) and D(n)), dominates in X(n), links are identified between the interference S(n) and a combination of the characteristics of the set {D(ni)} i ∈ DThis phenomenon is illustrated where there is an echo characterized by an overall delay of about 100ms. Around 34.9s, the window where the echo is alone, the standard attention module correctly identifies this overall delay of about 100ms. However, at the appearance of the useful signal s(t) around 35.25s, the standard attention model identifies relationships between the microphone and characteristics of D(ni) further away in time, while the overall delay characteristic of the echo has not changed.
[0097] In particular, the proposed attention mechanism is modified so as to integrate, in particular, an attention score filtering module (309) and an echo presence detector.
[0098] The attention score filtering module receives as input attention scores denoted a(n,i) from module (208), identical to enet. Unlike the classical attention mechanism described in equation [6], the expressions of the two signals are normalized, cf. normalization blocks (30a, 30b, 30c, 30d), after projection by the matricesW k andW q so that ||X p || = ||D p || = 1. This makes the attention mechanism insensitive to the scaling factor, e.g. the echo level in the microphone signal.
[0099] These attention scores provided as input to the filtering module are called "unweighted attention scores" in the remainder of this description. A set of unweighted attention scores can thus be obtained for each frame of the second signal.
[0100] The rest of the description of the implementation example considered focuses on the processing, by the attention score filtering module, of a set of unweighted attention scores a(n,i) obtained for a current frame of index n resulting from the decomposition or encoding of the signal captured by the microphone.
[0101] This processing can be repeated for each frame of the second signal.
[0102] The attention score filtering module relies on a weighting of the set of attention scores obtained with a set of weights, denoted h(n,i), derived from a combination of unweighted attention scores obtained for the temporal context D, thus obtaining weighted attention scores denoted a(n,i).
[0103] The detector of the presence of an echo, that is to say of the presence of a component ϕ(d(ni τ)) within the signal x(t) captured by the microphone, is based, directly or indirectly, on the weighted attention scores. It can be composed, for example, of a deep neural network whose architecture does not depend on the size of the temporal context, noted card(D), representing the number of possible delays considered. This detector can be used, in the context of echo cancellation, by a weighter of the characteristics D(ni) configured so as to transmit or not information depending on whether echo is present or not in the microphone or not.
[0104] Details of a possible implementation of the attention score filtering module (309) are now provided with reference to.
[0105] In a first step, the unweighted attention scores a(n,i) such as those provided by equation [6], are weighted by a series of weights h(n,i) where i ∈D. This weighting aims to concentrate the weighted attention scores a(n,i) on a particular index i τcorresponding to the overall delay τ and thus avoid dispersion of the weights over all possible delays i in the given temporal contextD. The calculation of the weights h(n,i) can be carried out from a combination of the unweighted attention scores a(n,i) in the following way: for each h(n,i), the combination involves all or part of the unweighted attention scores a(n,i) where i ∈D, as well as the scores of some adjacent frames a(n±k,:) with k ∈ N. Taking adjacent frames into account allows both smoothing to avoid rapid variations in the delay estimator, and also to correct the attention scores associated with the current frame of index n if they are very noisy. The number of adjacent frames can be determined experimentally. In general, a temporal context of one or two frames around the frame of index n is sufficient, i.e. k ∈ {1,2}.
[0106] From a computational point of view, this combination (402a, 402b) can be achieved by a series of 2D convolutional layers to obtain intermediate weights denoted h(n,i) where i ∈D. Unlike “fully-connected” layers, the use of convolutional layers avoids requiring an a priori on the card(D) dimension of the delay domain, which avoids retraining the network if, during inference, a temporal context D is used that is different from that of training.
[0107] An activation function (404) can be applied by a weight activation module to normalize the output weights. In the application to echo cancellation, it may be interesting to consider the weights h(n,i) associated with the current frame of index n as a probability on the indices i.
[0108] Also, it is possible to classically use the softmax function, which can be formalized according to the following equation: h(n,i) = eh(n,i) / ∑ i ∈ D e h(n,i)
[0109] It is also possible to use a L1-norm type normalization, which can be formalized according to the following equation.h(n,i) = |h(n,i)| / ∑ i ∈ D |h(n,i)|
[0110] Thanks to the normalization of the weights, the weights h(n,i) respect the relations 0 ≤ h(n,i) ≤ 1 and ∑ i ∈ D h(n,i) = 1 and can be interpreted as probabilities. After training, these weights h(n,i) can be interpreted as the probability, denoted P(H(n,i)), that the index i corresponds to the delay τ between the signal from the loudspeaker and the signal picked up by the microphone.
[0111] These probabilities can be used to weight (406) the attention scores a(n,i) of the current frame according to the following relationship, implemented by an attention score weighting operator. a(n,i) = h(n,i) a(n,i)
[0112] It may be interesting to reduce these new scores to a probability noted ã(n,i). To do this, it is possible to introduce a normalization type activation function, which can be the softmax function, or of type L1, or any other normalization. The following equations correspond respectively to softmax and L1 activation functions, implemented by a weighted score activation module (408).ã(n,i) = e a(n,i) / ∑ i ∈ D e a(n,i) ã(n,i) = |a(n,i)| / ∑ i ∈ D |a(n,i)|
[0113] During the learning process dedicated to echo cancellation, this weighting method facilitates the concentration of attention on the index iτ and allows the elimination of rapid variations from frame to frame, a phenomenon often encountered in standard attention models.
[0114] In terms of tuning, the weights of the neural network can be tuned in any way to best approximate the respective functions. Since the entire network is differentiable with respect to the parameters to be tuned, any optimization method based on gradient descent with respect to these parameters can be employed. The weights of the attention score filtering module can be learned in a supervised manner together with the other elements of the echo canceller. For example, a database can be built up comprising examples of signals x(t), their echoes e(t) picked up by a microphone after diffusion by a loudspeaker (real echoes recorded or synthesized by acoustic simulation) with different global delays τ, examples of useful signals s(t).
[0115] The weights of the entire echo canceller can then be learned in a supervised manner, by choosing a cost function L, such as the mean square error ℒ(s,ŝ) = 1 / T ∑ t (s(t) – ŝ(t)) 2 , or functions more suited to the perception of the human ear such as SI-SDR (Scale-Invariant Signal-to-Distortion-Rate) [LeRoux2019], or any other cost function using the ground truth s(t). It is also possible to train it in an unsupervised manner with methods such as MetricGAN for example [Fu2019].
[0116] The remainder of the description of the implementation example considered focuses on the processing, by the echo presence detector, of the sets of weighted attention scores a(n,i) or ã(n,i) obtained for a plurality of frames resulting from the decomposition or encoding of the signal captured by the microphone.
[0117] This processing provides, from the weighted attention scores, to deduce a probability of echo presence, noted P(E(n)), at the microphone level. The calculation of P(E(n)) consists of combining the scores ā(n±k,i) or ã(n±k,i) where i ∈Det where k is a natural integer.
[0118] From a computational point of view, this combination can be achieved by the architecture presented in.
[0119] In particular, a first combination (410a, 410b can be carried out by a series of 1D convolutional layers) in the delay axis i, with, at each layer, a sub-sampling of the delay axis.
[0120] The residue of the delay dimension can be removed (412) to make the rest of the neural network independent of the size of this dimension. To do this, it is possible to apply a pooling layer to this dimension, which can be of the average, median, or maximum type, for example. By nature, the periods of presence or absence of echo are stable over relatively long periods compared to the duration of a frame. Also, the use (414) of a recursive neuron of the LSTM or GRU type, by smoothing the predictions from the convolutional layers by taking into account the short- and long-term contexts, makes it possible to avoid overly rapid variations, which are not representative of real conditions.
[0121] Finally, to be analogous to a probability, it is possible to choose an output activation function of the recursive neuron configured to provide an output value between 0 and 1. This probability P(E(n)) can be used to weight once again the weighted attention scores a(n,i) or ã(n,i) and thus produce the final estimate of the overall delay between the echo and the useful signal captured by the microphone. This last weighting (416) can be formalized, in the case where the weighted attention scores are the normalized scores ã(n,i), by the following equation implemented by a weighted attention score weighting operator.â(n,i) = P(E(n)) . ã(n,i)
[0122] Unlike standard attention scores, the scores â(n,i) do not sum to 1 and can all be zero. This allows, in the absence of echo where P(E(n) ≪ 1, to indicate to the echo canceller that the loudspeaker signal should not be taken into account when calculating the maskm(n) and thus to avoid degradation of the useful signal.
[0123] Steps 210a, 210b, 210c and 212 illustrated in and as described with reference to are implemented with the weighted attention scores from the filtering module 309. Thus a linear combination of the characteristics of the first signal with the weighted attention scores is carried out to obtain an estimate of the realigned characteristics.
[0124] In the examples in enet, the same examples as in enet are illustrated, but using a modified attention mechanism according to the proposed technique rather than a standard attention mechanism. First, we notice that when ϕ(.)=0 then the attention score is this time close to zero (608), i.e., there is nothing to realign (606), resulting inD r (n) close to zero, where the classic attention module tried at all costs to realign the first signal with the second, even if they have nothing in common (506, 508). Then, we observe that the attention score (808) is much more stable over time, essentially conveying the information of delay (around 100 ms) and presence or absence of echo in the microphone, as expected.
[0125] The above description presents an attention score filter which, combined with a standard attention mechanism, forms a modified attention module, as well as an acoustic echo presence detector.
[0126] The proposed attention score filter or the modified attention module integrating such a filter can be integrated into any application requiring signal rephasing.
[0127] The acoustic echo presence detector is a particular example of a detector for a particular application. Generally, the attention score filter can be associated with a module for detecting an expression of a first signal in a second signal decomposed into a plurality of frames.
[0128] In the particular case where the detection module is an acoustic echo presence detector, the first signal is the signal emitted by the loudspeaker, the second signal is the signal picked up by the microphone and the expression of the first signal and the expression of the first signal in the second signal represents the acoustic echo.
[0129] Besides its application in acoustic echo cancellation, the proposed technique also finds its use in determining the origin of two signals captured separately. For example, the estimated overall delay between the expression of a first signal and that of a second signal can serve as input data for a triangulation mechanism, aiming to locate a common spatial and / or temporal origin of the two signals. This approach is particularly relevant in fields such as the registration of two electrocardiograms obtained via sensors placed on different parts of the body, medical diagnosis by ultrasound, geolocation and structural diagnosis by seismographic techniques.
[0130] Another possible use of the proposed technique concerns signal synchronization. The estimated global delay can be used to temporally coordinate two signals, such as in the registration of two videos of the same scene captured by separate cameras. In this case, the dimensions of the convolutional layers of the registration module must be adapted to the two-dimensional data. Moreover, this method finds practical applications in the multimedia domain, such as lip synchronization in video conferencing or synchronization of audio tracks in film post-production, thus ensuring increased consistency and quality of the user experience.
[0131] Although the proposed technique has been described in detail with reference to certain preferred embodiments, various modifications and variations may be made without departing from the scope of the invention as defined in the appended claims. For example, the described procedures may be performed in a different order, elements may be added or deleted, and features of different embodiments may be combined as appropriate.
[0132] represents a device (900), or a system, for estimating delay between a first signal and a second signal.
[0133] It comprises in particular: an output module (902), a reception module (904), a processor (906) or signal processing unit, and a memory (908) or data storage unit adapted to store instructions of a computer program which, when the program is implemented by the processor, lead to implementing the proposed technique.
[0134] Such a device or system may be embedded in a terminal equipped with a loudspeaker configured to emit the first signal and a microphone configured to receive the second signal. In such a terminal, the loudspeaker is then connected to the output module and the microphone is connected to the reception module to allow, for example, the implementation of an echo cancellation mechanism according to an exemplary embodiment.
[0135] The device or system may not include an output module and the receiving module (904) may be configured to receive both the first signal and the second signal. The receiving module may, for example, be connected to a first sensor configured to receive the first signal and to a second sensor configured to receive the second signal, the first sensor and the second sensor being placed at separate locations. Such a configuration may, for example, allow the implementation of a mechanism for determining an origin of the first signal and the second signal according to an exemplary embodiment. Alternatively, the receiving module may be connected to at least one sensor configured to receive the first signal and the second signal separately, for example at separate times.Such a configuration may allow, for example, the implementation of a method of synchronizing the first signal and the second signal according to an exemplary embodiment.
Claims
A method of temporal alignment between characteristics of a first signal and characteristics of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the method comprising, for a current frame of the second signal:obtaining (208) a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, andobtaining (309), by a trained filtering module, a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally neighboring the current frame,weighting (309) the attention scores with the weights obtained to produce a set of weighted attention scores;a linear combination (212) of the features of the first signal with the weighted attention scores to obtain an estimate of the realigned features.; The method of claim 1, wherein the characteristics of the two signals are normalized so that their norm is equal to 1. The method of claim 1, wherein the resulting set of weighted attention scores is normalized to correspond to a set of probabilities. The method of claim 1, further comprising detecting features of the first signal in the second signal by determining a probability of presence of features of the first signal in the second signal based on processing the sets of weighted attention scores comprising implementing one-dimensional convolutional layers acting on a time axis of delays, each convolutional layer being configured to perform a sub-sampling of the time axis of delays and implementing a pooling layer to derive the probability of presence therefrom. Method according to claim 4, further comprising a step of smoothing the probabilities obtained by implementing a recursive neuron. A method of cancelling an echo represented by characteristics of a first signal within a second signal, the method comprising aligning the characteristics of the first signal with the characteristics of the second signal according to the method of one of claims 1 to 5. A method of determining an origin of a first signal and a second signal, the first signal and the second signal being sensed separately, based on a delay between characteristics of the first signal and characteristics of the second signal, the method comprising an alignment between the characteristics of the first signal and the characteristics of the second signal according to the method of one of claims 1 to 5. A system for aligning features of a first signal with features of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the system comprising the following modules applied to a current frame of the second signal: a module for obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, a trained filtering module (309) capable of - obtaining a set of weights (402a, 402b, 404), derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally adjacent to the current frame, and - weighting (406) attention scores with the weights obtained to produce a set of weighted attention scores;a combination module (212) capable of linearly combining features of the first signal with the weighted attention scores to obtain an estimate of the realigned features.; Computer program comprising instructions which, when the program is implemented by a processor, lead to implementing the method according to one of claims 1 to 7.
Citation Information
Patent Citations
Voice chip and electronic equipment
CN109584896A
Temporal alignment of signals using attention
WO2023219751A1