Attention mechanism for delay estimation and alignment of two signals

The attention mechanism-based method effectively addresses the challenge of estimating signal delays in variable environments by concentrating weighted attention scores around a specific delay, enhancing echo cancellation and signal alignment.

FR3157646A1Inactive Publication Date: 2025-06-27ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2023014939
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for estimating the delay between two signals, particularly in contexts with variable environments and intermittent or absent interference, are insufficient in handling complex temporal variations and echo absence scenarios.

Method used

A method and system that utilize an attention mechanism to estimate the delay between two signals by obtaining unweighted attention scores and deriving weights from a temporal context, then weighting the attention scores to concentrate around a specific delay, thereby providing an accurate estimate of the overall delay.

Benefits of technology

The proposed method achieves accurate and reliable delay estimation, particularly in contexts where the delay remains relatively constant, and improves echo cancellation and signal alignment by efficiently handling scenarios with absent or complex echo patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000022_0001
    Figure 00000022_0001
  • Figure 00000023_0000
    Figure 00000023_0000
Patent Text Reader

Abstract

A method is proposed for estimating a delay between an expression of a first signal and an expression of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the method comprising, for a current frame of the second signal:obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, andobtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally close to the current frame,a weighting (406) of the attention scores with the weights obtained to produce a set of weighted attention scores concentrated around a specific delay corresponding to an estimated overall delay between the expression of the first signal and that of the second signal. Abstract figure: Figure 4,
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Attention mechanism for estimating a delay and aligning two signals Technical field

[0001] The present disclosure is in the field of information and communications technologies, and more specifically in the field of signal processing.

[0002] It relates to a method and a system for estimating a delay between an expression of a first signal and an expression of a second signal, a computer program for implementing such a method and a non-transitory recording medium readable by a computer on which such a program is recorded. Prior art

[0003] Accurately estimating the delay between two potentially coupled signals is a problem common to several technical fields. Current solutions are often designed for specific use cases.

[0004] For example, current voice communication systems, especially those used in hands-free environments, are frequently confronted with the problem of echo. This echo is mainly due to impedance differences in the wired connections and to acoustic or solid couplings between the loudspeakers and the microphones. Traditional echo cancellation techniques involve the use of adaptive but also non-linear filters operating in the frequency domain.

[0005] Although current methods, including those using neural networks for time-frequency mask estimation, have improved the performance of echo cancellation, they remain insufficient in certain scenarios. In particular, they struggle to effectively handle situations where the echo is absent, has complex temporal variations, or when interference is added to the echo.

[0006] There is a need for a robust method capable of accurately estimating a delay between two signals in a variety of contexts, particularly in variable environments where interference between these two signals may be intermittent or absent. Summary

[0007] The invention aims to improve the situation.

[0008] According to one aspect, there is proposed a method for estimating a delay between an expression of a first signal and an expression of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the method comprising, for a current frame of the second signal: obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, and obtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally adjacent to the current frame, a weighting of the attention scores with the obtained weights to produce a set of weighted attention scores concentrated around a specific delay corresponding to an estimated overall delay between the expression of the first signal and that of the second signal.

[0009] According to one aspect, there is proposed a system for estimating a delay between an expression of a first signal and an expression of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the system comprising the following modules applied to a current frame of the second signal: a module for obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, a module for obtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally adjacent to the current frame, and an attention score weighting module with the weights obtained by the weight set obtaining module to produce a set of weighted attention scores concentrated around a specific delay corresponding to an estimated overall delay between the expression of the first signal and that of the second signal.

[0010] A technical advantage of the proposed method and system lies in their ability to provide an accurate and reliable estimate of the overall delay between the expression of the first signal and the expression of the second signal. This is particularly beneficial in contexts where the overall delay between the expressions of the first signal and the second signal remains relatively constant from one frame to another. In such contexts, a constancy or limited variation of the sets of weighted attention scores from frame to frame indicates a stability of the estimated overall delay.

[0011] In one example, the expressions of the two signals are normalized so that their norm is equal to 1. This makes the system independent of the difference in scale factor between the two signals.

[0012] In one example, the resulting set of weighted attention scores is normalized to way to correspond to a set of probabilities.

[0013] This normalization makes it possible to provide consistent information over time, from one frame to another, for processing modules located downstream. Also proposed, according to one aspect, is a method for detecting an expression of a first signal in a second signal decomposed into a plurality of frames, the method comprising an estimation of a delay between the expression of the first signal and that of the second signal according to the proposed estimation method and further comprising a determination of a probability of presence of the expression of the first signal in the second signal on the basis of a processing of the sets of weighted attention scores.

[0014] Detection of the expression of a first signal within a second signal is improved by processing weighted attention score sets, thereby significantly reducing false detections, resulting in more robust alignment.

[0015] In one example, determining the probability of presence includes implementing one-dimensional convolutional layers acting on a delay axis, each convolutional layer configured to downsample the delay axis.

[0016] This implementation offers the advantage of a fine and detailed analysis of the signals, allowing precise localization of the expression of the first signal within the second signal and reliable detection even in the presence of complex or noisy signals.

[0017] In one example, the combination of the sets of weighted attention scores comprises, after the implementation of the convolutional layers, an implementation of a pooling layer configured to make the probability of presence of the expression of the first signal within the second independent of the number of frames forming the given temporal context.

[0018] This makes it possible to change the size of the temporal context considered without having to re-train the neural network. Also proposed, according to one aspect, is a method for canceling an echo represented by an expression of a first signal within a second signal, the method comprising an estimation of a delay between the expression of the first signal and an expression of the second signal according to the proposed estimation method or a detection of the expression of the first signal in the second signal according to the proposed detection method.

[0019] Effective echo cancellation requires having an accurate estimate of the delay between the expression of the first signal to be canceled and the expression of a second signal containing a useful signal to be retained. This accurate estimate is provided by the use of weighted attention scores. Furthermore, the implementation of the method The proposed detection method allows, thanks to the processing of sets of weighted attention scores, to efficiently detect the presence or absence of echo and, consequently, to preserve the quality of a useful signal by avoiding processing a non-existent echo.

[0020] Also provided, according to one aspect, is a method for determining an origin of a first signal and a second signal, the first signal and the second signal being captured distinctly, based on a delay between an expression of the first signal and an expression of the second signal, the method comprising an estimation of the delay according to the proposed estimation method or a detection of the expression of the first signal in the second signal according to the proposed detection method.

[0021] This method offers the advantage of accurately tracing the source of the signals in the event of multiple transmission or reflection paths, which is essential in applications such as medical diagnosis by ultrasound, geolocation and structural diagnosis by seismography techniques.

[0022] Also provided, according to one aspect, is a method for synchronizing a first signal and a second signal, the method comprising estimating a delay between an expression of the first signal and an expression of the second signal according to the proposed estimation method or detecting the expression of the first signal in the second signal according to the proposed detection method.

[0023] This method offers the advantage of precise coordination between signals, essential in multimedia applications such as lip synchronization in video conferencing or synchronization of audio tracks in film post-production, where even a minimal time shift can be noticeable and annoying.

[0024] Also provided, in one aspect, is a computer program comprising instructions which, when the program is implemented by a processor, result in implementing one of the methods discussed.

[0025] Also provided in one aspect is a non-transitory computer-readable recording medium having recorded thereon a program for implementing one of the methods discussed when that program is executed by a processor. Brief Description of the Drawings

[0026] Other characteristics, details and advantages will appear on reading the detailed description below, and on analyzing the attached drawings, in which: Fig.l

[0027] [Fig.l] represents an acoustic echo cancellation algorithm according to the state of the art. Fig. 2

[0028] [Fig.2] represents an attention model applied to acoustic echo cancellation according to the state of the art. Fig. 3

[0029] [Fig.3] represents an attention model applied to acoustic echo cancellation according to an example of realization. Fig. 4

[0030] [Fig.4] represents an attention score filtering module and a detector of presence of echo suitable for the attention model of [Fig.3]. Fig. 5

[0031] [Fig.5] represents an example of the result of a time alignment of two signals in the absence of echo, the time alignment using a delay determined by application of the model of [Fig.2]. Fig. 6

[0032] [Fig.6] represents an example of the result of a time alignment of two signals in the absence of echo, the time alignment using a delay determined by application of the model of [Fig.3]. Fig. 7

[0033] [Fig.7] represents an example of the result of a time alignment of two signals in the presence of interference and echo, the time alignment using a delay determined by application of the model of [Fig.2]. Fig. 8

[0034] [Fig.8] represents an example of the result of a time alignment of two signals in the presence of interference and echo, the time alignment using a delay determined by application of the model of [Fig.3]. Fig. 9

[0035] [Fig.9] represents an example of a device for estimating the delay between a first signal and a second signal. Description of the embodiments

[0036] In the following description, identical reference numerals designate identical elements or elements having similar functions.

[0037] The following description mentions the following references.

[0038] [Indenbom2022] Evgenii Indenbom, Nicolae-Câtâlin Ristea, Ando Saabas, Tanel Pârnamaa, and Jegor Guzvin, “Deep model with built-in self-attention alignment for acoustic echo cancellation,” Aug. 2022, arXiv:2208.11308 [es, eess].

[0039] [LeRoux2019] JL Roux, S. Wisdom, H. Erdogan, and JR Hershey, “Sdr - half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.

[0040] [Fu2019] Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, Shou-De Lin, “MetricGAN: Ge- nerative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement”, International Conference on Machine Leaming, 2019.

[0041] The following description refers to known technologies in the field of signal processing, in particular known artificial intelligence technologies.

[0042] A presentation of a case study in which such technologies are used for echo cancellation is now provided to facilitate understanding of the proposed technique.

[0043] During a voice communication, particularly hands-free, an echo phenomenon may appear. Several causes are at the origin of this echo. Wired hybrid connections cause differences in impedances, sources of audible reflections such as an echo with a certain delay. On the terminal side, couplings exist between the loudspeaker and the microphone. These couplings can be of an acoustic nature, particularly in the case of hands-free communications: the reverberation or room effect, due to the walls, create multiple reflections which will be picked up by the microphone. There are also solid couplings within the terminal due to the vibration of certain parts, vibrations picked up by the microphone.Due to these couplings, the voice of the remote interlocutor broadcast by the loudspeaker is picked up by the microphone with a certain delay and sent back to the remote party, creating an echo phenomenon which degrades the communication and impacts the intelligibility of the local interlocutor. Also, telephone systems integrate echo cancellers at different points in the chain in order to eliminate this artifact.

[0044] [Fig-1] illustrates an echo canceller (102) configured to cancel, or at least attenuate, an echo representing an expression of a signal emitted by a loudspeaker (104) within a signal picked up by a microphone (106).

[0045] The following notations are introduced: d(t) denotes the signal emitted by the loudspeaker, ¢(.) denotes the coupling between the loudspeaker and the microphone, e(t) denotes the echo picked up by the microphone, s(t) denotes the useful signal, and x(t) denotes the signal picked up by the microphone.

[0046] x(t) is given by the following additive relation, denoted [1]. x(t)=s(t)+e(t)=s(t)+0(d(t))

[0047] The coupling ¢)(.) between the loudspeaker and the microphone is characterized by: linear transformations representative of the frequency response of the transducers (loudspeaker / microphone) and the reverberation of the room, non-linear transformations, such as saturations which can occur at the level of the transducers, and an overall delay, noted r, which represents the sum of the delay of the terminal's acquisition-restitution chain (jitter buffers, digital conversions- analog) and the acoustic delay linked to the propagation of sound in the air.

[0048] Echo cancellation techniques typically use a serial combination of an adaptive filter seeking to estimate the linear transformation and a non-linear filter seeking to remove the residual echo at the output of the adaptive filter. For computational and performance reasons, the processing is generally carried out in the frequency domain: this makes it possible, on the one hand, to design more robust and faster adaptive filters, and on the other hand, to develop non-linear filters in the form of time-frequency masks which guarantee a better compromise between the suppression of the residual echo and voice quality.

[0049] Thus, each signal denoted z(t), where z(t) designates one of the signals introduced previously, in particular d(t) and x(t), is decomposed into frames - adjacent or not - before applying a short-term Fourier transform to produce a frequency vector denoted Z(n) e CF, that is to say D(n) or X(n) respectively, where n is the index of the frame and F is the number of frequencies considered. Beforehand, an apodization window denoted w G RN, where N is a natural integer, is applied to each signal frame in order to control the effects of the windowing. The apodization window can for example be a Hann, Hamming, or Blackman window.

[0050] The application of the apodization window can be formalized by the following equation: Z(n) = J(z(n)0w) where z(n) is a vector of size N representing the frame of index n of the signal z(t), denotes a short-term Fourier transform, and O denotes a Hadamard product, or term-by-term product, of two vectors or matrices.

[0051] The superposition rate of two successive frames, noted (NP) / N, is fixed by the number P. P is a natural integer between 1 and N. When P=N, the superposition rate is equal to zero, indicating that the frames are adjacent.

[0052] The terms of the vector z(n) are noted z(n) =[z(n*P) z(n*(Pl)) ••• z(n*(P-N+l))].

[0053] Under this formalism, the non-linear module generally has the following functions: calculating a time-frequency mask m(n) G RF applying this mask to the output signal of the adaptive filter to estimate the useful signal in the time-frequency domain, noted S(n) and reconstruct the time version of the estimated useful signal, noted s(t), by inverse Fourier transform and by application of known methods of the “overlap-and-add” or “overlap-and-save” type.

[0054] In recent years, the use of deep neural networks to estimate the mask m(n) has shown very interesting performances [Indenbom2022], in any case much superior to signal processing techniques based on the estimation of Wiener filters from signal models (probabilistic models) and the coupling function ¢(.). The very latest approaches are sufficiently robust to no longer use an adaptive filter.

[0055] These approaches provide for applying the mask directly to the microphone signal. This application can be formalized by the following equation, denoted [2]. S(n)=m(n)oX(n)

[0056] In the latter case, the transform can be learned by the neural network. The architecture of the neural network then includes an encoding step, denoted T(.) implemented by a series of layers of neurons whose role is to extract characteristics of the signal considered, namely the signal emitted by the loudspeaker and the signal captured by the microphone, respectively.

[0057] The encoding carried out by the neural network can be formalized by the following equation, considering Q as being the set of indices of past or future frames to be provided as input to the encoder: Z(n) = r({z(nq)]qeQ)

[0058] In the remainder of this description, the use of the notation Z(n) refers to a frequency vector resulting, indifferently, from a short-term Fourier transform or from an encoding step carried out by the neural network.

[0059] To estimate the mask m(n) of the frame of index n, the neural network is fed with the characteristics of X(n) and with the characteristics of D(ni).

[0060] The index i here delimits the relevant temporal context, noted D, for the frame of index n.

[0061] This temporal context is defined by a set of frames encompassing both past and future frames relative to the frame of index n, expressed in the form {,-n_future < i < n_past}. In other words, for each frame of index n considered, the neural network takes into account a set of neighboring frames, extending from n_past frames in the past to n_future frames in the future.

[0062] The estimation of the mask is formalized by the following equation noted [3]. m(n) = W(X(n),D(n))

[0063] W(.) is a mask generation function. Its implementation can take the form of an algorithm or all or part of a neural network.

[0064] Conventionally, due to the overall delay r, echo cancellation systems integrate a time realignment module which consists of aligning the loudspeaker signal with the microphone signal by compensating for this delay, which allows better conditioning of the different estimators (adaptive filter and / or echo suppressor). Thus, the input of the echo cancellation system becomes d(tr), where r is an estimate of the delay r, and all the characteristics of the loudspeaker are calculated from this rephased signal. In the case of neural networks, rather than explicitly estimating the delay r, the network integrates in its architecture a realignment module which exploits a set, noted {D(ni)]i£D, of the characteristics of D(ni). The realignment by the realignment module is formalized by the following equation noted [4]. m(n) = W(X(n),{D(ni)U)

[0065] The temporal context D is to be determined according to different constraints. The maximum past n_past to be taken into account depends on the overall delay r and the duration of the reverberation of the room. In practice, n_past can take a value corresponding to several hundred milliseconds. As for the value of n_future which fixes the quantity of future that one allows to take, it depends on the causality constraint that one imposes: if one wants a processing without delay, as for example for telephony applications, one can fix n_future = 0; in other situations where the strict causality constraint is relaxed, one can allow to take one or more frames of the future to better estimate the mask m(n) of the current frame of index n.

[0066] In [Indenbom2022], the authors use an attention mechanism to achieve this realignment. The principle of the attention mechanism applied to echo cancellation consists of linearly combining the set {D(ni)}i£D of the characteristics of D(ni) with weights noted a(n,i) so as to obtain an estimator of the realigned characteristics, noted Dr(n), according to the following equation noted [5]. Dr(n)=EieD a(n,i).D(ni)

[0067] The weights a(n,i), called attention scores and calculated at each frame of index n, are obtained using a standard attention model which can be formalized according to the following equation, noted [6]. a(n) = softmax((l / 'Vb) Xp(n)T [Dp(n+n_future), ... , Dp(n-n_past)])

[0068] Where a(n) is the vector containing the a(n,i), Xp(n) = WqT. X(n) and Dp(n) = WqT. D(n) where Wq G Rcxb and Wk G Rvxb are matrices called "query" and "key" with c and v the dimensions of X(n) and D(n), respectively, and b the size of the space into which the expressions of the signals are projected.

[0069] The aforementioned attention mechanism that can be implemented within an echo canceller is illustrated in [Fig.2].

[0070] A first matrix multiplication operator (202) performs the product of the vector X(n)T and the “query” matrix Wq.

[0071] For each vector D(ni), a second corresponding matrix multiplication operator (204a, 204b, 204c) performs the product of the “key” matrix WkT and the aforementioned vector D(ni).

[0072] For each vector D(ni), a third corresponding matrix multiplication operator (206a, 206b, 206c) performs the product of the output vector WkT D(ni) of the second corresponding operator and the output vector X(n)T Wq of the first operator.

[0073] An attention module (208) based on a neural network and comprising a softmax type activation function receives as input the outputs of the third operators and calculates as output a set of attention scores a(n,i), i.e. an attention score for each vector D(ni).

[0074] For each vector D(ni), a corresponding multiplication operator (210a, 210b, 210c) performs the product of the corresponding attention score a(n,i) and the aforementioned vector D(ni).

[0075] An addition operator (212) sums the outputs of the multiplication operators (210a, 210b, 210c) and thus obtains the estimator of the realigned characteristics according to equation [5].

[0076] In the context of time alignment, we expect a(n,i) to be maximum, i.e. close to 1, around the index noted iT corresponding to the overall delay r, and to take a low value, i.e. close to 0, elsewhere.

[0077] The index iT is calculated as follows, as being the integer part of the quotient of r by P: iT=lr / PI

[0078] In light of the case study as provided above, the proposed technique is now presented in the context of an application to echo cancellation.

[0079] The proposed technique is illustrated in [Fig.3] and [Fig.4] and is to be compared with [Fig.2],

[0080] The results obtained by applying the proposed technique to test scenarios are illustrated in [Fig.6] and [Fig.8], and can be compared to the results obtained by applying an attention mechanism according to [Fig.2], illustrated in [Fig.5] and [Fig.7],

[0081] [Fig.5], [Fig.. 6], [Fig.7] and [Fig.8] each comprise four representations, know : a representation (502, 602, 702, 802) of the characteristics of the signal emitted by the loudspeaker as a function of time, a representation (504, 604, 704, 804) of the characteristics of the signal picked up by the microphone as a function of time, a representation (506, 606, 706, 806) of the characteristics of the signal emitted by the loudspeaker synchronized by the system with the signal picked up by the microphone as a function of time, and a representation (802, 806) of the attention scores as a function of time in [Fig.5] and [Fig.7] and a representation (804, 808) of the weighted attention scores as a function of time in [Fig.6] and [Fig.8].

[0082] The proposed technique provides in this context a time alignment solution robust to interference, and to the presence or absence of echo in the signal picked up by the microphone. The proposed technique makes it possible in particular to obtain, unlike the behavior observed by a classic attention module, scores which reflect desired characteristics for a delay estimator between the signal picked up by a microphone and the signal emitted by a loudspeaker and an expression of which is likely to be present in the signal picked up by the microphone.

[0083] In particular, according to the proposed technique, each frame of index n is associated with a set of scores which can be represented by a vector noted â(n) = [â(n,n_future), ... , â(n,n_past)].

[0084] This set of scores has values ​​concentrated around a particular delay corresponding to the overall delay between the loudspeaker and the microphone.

[0085] In the continuous presence of echo, the set of scores associated with each frame is stable over time, i.e. from frame to frame, including in the presence of useful signal s(t) captured by the microphone.

[0086] Furthermore, unlike a standard attention mechanism, the proposed technique makes it possible to efficiently model echo-free scenarios. In particular, the vector â(n) can be designed to be representative not only of a delay estimate, but also of a presence or absence of coupling between the signal emitted by the loudspeaker and the signal picked up by the microphone. For example, the vector â(n) can be designed to have a variable norm depending on the presence or absence of such coupling. For example, the vector â(n) can be designed to have a norm close to zero when the coupling between the loudspeaker signal and the microphone signal is zero (¢(.)=0).

[0087] Indeed, the standard attention model is originally designed for application to natural language processing and has properties that are not ideally suited to application to the temporal realignment of audio signals in the context of echo cancellation.

[0088] The standard attention model constrains the attention scores to sum to one, through the use of the softmax function in equation [6]. This constraint is problematic because, in the context of echo cancellation, the coupling between the signal emitted by the loudspeaker and the signal picked up by the microphone can be zero, i.e. ¢(.)=0. In the absence of echo, it is not desirable for the model to identify a correlation between the microphone and the loudspeaker. Indeed, the presence of non-zero values ​​a(n,i) amounts to feeding the echo canceller with characteristics from the signal coming from the loudspeaker, which indicates to the echo canceller the presence of echo and leads in practice to attenuate, and therefore potentially degrade, a local speech signal s(t)=x(t) when no echo is present.

[0089] This case is illustrated in [Fig.5] where the loudspeaker emits a signal which is not present at the microphone level because the coupling is zero.

[0090] It is observed that the standard attention mechanism, represented in (508), still identifies links between X(n) and characteristics of the set {D(ni)]i£D of the characteristics of D(ni). In this case, the identified links do not reflect a delay estimate but simply indicate to the echo canceller that it is still necessary to take into account certain characteristics D(ni) to calculate the mask m(n).

[0091] In an application of the standard attention model to natural language processing, the attention score can vary greatly from one moment to the next. However, in an application to echo cancellation, the delay rarely varies and it is desirable that a(n,i).D (ni) varies little from frame to frame. However, it is observed in practice that with standard attention models, if the interference, expressed in the form of a frequency vector S(n) (like X(n) and D(n)), dominates in X(n), links are identified between the interference S(n) and a combination of the characteristics of the set {D (ni)]i&D. This phenomenon is illustrated in [Fig.7] where there is an echo characterized by an overall delay of approximately 100ms. Around 34.9 s, the window where the echo is alone, the standard attention module correctly identifies this overall delay of around 100 ms.However, at the appearance of the useful signal s(t) around 35.25s, the standard attention model identifies relationships between the microphone and features of D(ni) more distant in time, while the overall delay characteristic of the echo has not evolved.

[0092] In particular, the attention mechanism proposed in [Fig.2] is modified so as to integrate, in particular, a filtering module (309) of attention scores and an echo presence detector.

[0093] The attention score filtering module receives as input attention scores denoted a(n,i) from the module (208), identical in [Fig.2] and [Fig.3]. Unlike the classic attention mechanism described in [Fig.2] and equation [6], the expressions of the two signals are normalized, cf. normalization blocks (30a, 30b, 30c, 30d), after projection by the matrices Wk and Wq so that IIXPII = IIDPII = 1. This makes it possible to make the attention mechanism insensitive to the scale factor, eg the echo level in the microphone signal.

[0094] These attention scores provided as input to the filtering module are called “unweighted attention scores” in the remainder of this description. A set of unweighted attention scores can thus be obtained for each frame of the second signal.

[0095] The remainder of the description of the implementation example considered focuses on the processing, by the attention score filtering module, of a set of unweighted attention scores a(n,i) obtained for a current frame of index n resulting from the decomposition or encoding of the signal captured by the microphone.

[0096] This processing can be repeated for each frame of the second signal.

[0097] The attention score filtering module is based on a weighting of the set of attention scores obtained with a set of weights, denoted h(n,i), derived from a combination of unweighted attention scores obtained for the temporal context D, thus obtaining weighted attention scores denoted a(n,i).

[0098] The echo presence detector, i.e. the presence of a component 0(d(n-iT)) within the signal x(t) captured by the microphone, is based, directly or indirectly, on the weighted attention scores. It can be, for example, composed of a deep neural network whose architecture does not depend on the size of the temporal context, denoted card(D), representing the number of possible delays considered. This detector can be used, in the context of echo cancellation, by a weighter of the characteristics D(ni) configured so as to transmit or not information depending on whether echo is present or not in the microphone or not.

[0099] Details of a possible implementation of the attention score filtering module are now provided with reference to [Fig.4].

[0100] In a first step, the unweighted attention scores a(n,i) such as those provided by equation [6], are weighted by a series of weights h(n,i) where ie D. This weighting aims to concentrate the weighted attention scores a(n,i) on a particular index iT corresponding to the overall delay r and thus avoid the dispersion of the weights over all possible delays i in the given temporal context D. The calculation of the weights h(n,i) can be carried out from a combination of the unweighted attention scores a(n,i) in the following manner: for each h(n,i), the combination involves all or part of the unweighted attention scores a(n,i) where ie D, as well as the scores of some adjacent frames a(n±k,:) with ke N.Considering adjacent frames allows both smoothing to avoid rapid variations in the delay estimator, and also correcting the attention scores associated with the current frame of index n if they are very noisy. The number of adjacent frames can be determined experimentally. In general, a temporal context of one or two frames around the frame of index n is sufficient, i.e. ke {1,2}.

[0101] From a computational point of view, this combination (402a, 402b) can be carried out by a series of 2D convolutional layers to obtain intermediate weights noted h(n,i) where ie D. Unlike “fully-connected” type layers, the use of convolutional layers avoids requiring an a priori on the card(D) dimension of the delay domain, which avoids retraining the network if, during inference, a temporal context D different from that of the learning is used.

[0102] An activation function (404) may be applied by a weight activation module in order to normalize the output weights. In the application to echo cancellation, it may be interesting to consider the weights h(n,i) associated with the current frame of index n as a probability on the indices i.

[0103] Also, it is possible to classically use the softmax function, which can be formalized according to the following equation. h(n,i) = e^114' / eh(n4)

[0104] It is also possible to use a normalization of the Ll norm type, which can be formalized according to the following equation. h(n,i) = lh(n,i)l / Ei£D lh(n,i)l

[0105] Thanks to the normalization of the weights, the weights h(n,i) respect the relations 0 < h(n,i) < 1 and h(n,i) = 1 and can be interpreted as probabilities. After learning, these weights h(n,i) can be interpreted as the probability, noted P(H(n,i)), that the index i corresponds to the delay r between the signal coming from the loudspeaker and the signal picked up by the microphone.

[0106] These probabilities can be used to weight (406) the attention scores a(n,i) of the current frame according to the following relationship, implemented by an attention score weighting operator. a(n,i) = h(n,i) a(n,i)

[0107] It may be interesting to reduce these new scores to a probability noted â(n,i). For this, it is possible to introduce a normalization type activation function, which can be the softmax function, or of type Ll, or any other normalization. The following equations correspond respectively to softmax and Ll activation functions, implemented by a weighted score activation module (408). â(n,i) = ea(n4) / EieD ea(n4) â(n,i) = la(n,i)l / EieD la(n,i)l

[0108] During the learning process dedicated to echo cancellation, this weighting method facilitates the concentration of attention on the ir index and makes it possible to eliminate rapid variations from frame to frame, a phenomenon often encountered in standard attention models.

[0109] In terms of tuning, the weights of the neural network can be tuned in any way so as to best approximate the respective functions. Since the entire network is differentiable with respect to the parameters to be tuned, any optimization method based on gradient descent with respect to these parameters can be used. The weights of the attention score filtering module can be learned in a supervised manner together with the other elements of the echo canceller. For example, a database can be built up comprising examples of signals x(t), their echoes e(t) picked up by a microphone after diffusion by a loudspeaker (real echoes recorded or synthesized by acoustic simulation) comprising different overall delays r, examples of useful signals s(t).

[0110] The weights of the entire echo canceller can then be learned in a supervised manner, by choosing a cost function L, such as the mean square error J^s,s) = 1 / T Et (s(t) - s(t))2, or functions more suited to the perception of the human ear such as the SLSDR (Scale-Invariant Signal-to-Distortion-Rate) [LeRoux2019], or any other cost function using the ground truth s(t). It is also possible to train it in an unsupervised manner with methods such as MetricGAN for example [Fu2019].

[0111] The remainder of the description of the implementation example considered focuses on the processing, by the echo presence detector, of the sets of weighted attention scores a(n,i) or â(n,i) obtained for a plurality of frames resulting from the decomposition or encoding of the signal captured by the microphone.

[0112] This processing provides, from the weighted attention scores, to deduce a probability of echo presence, noted P(E(n)), at the microphone level. The calculation of P(E(n)) consists of combining the scores â(n±k,i) or â(n±k,i) where ie D and where k is a natural integer.

[0113] From a computational point of view, this combination can be achieved by the architecture presented in [Fig.4].

[0114] In particular, a first combination (410a, 410b can be carried out by a series of 1D convolutional layers) in the delay axis i, with, at each layer, a sub-sampling of the delay axis.

[0115] The residue of the delay dimension can be removed (412) to make the sequence of the neural network independent of the size of this dimension. For this, it is possible to apply to this dimension a grouping layer which can be of the average, median or maximum type, for example. By nature, the periods of presence or absence of echo are stable over relatively long periods compared to the duration of a frame. Also, the use (414) of a recursive neuron of the LSTM or GRU type, by smoothing the predictions from the convolutional layers by taking into account the short and long-term contexts, makes it possible to avoid overly rapid variations, which are not representative of real conditions.

[0116] Finally, to be analogous to a probability, it is possible to choose an output activation function of the recursive neuron configured to provide an output value between 0 and 1. This probability P(E(n)) can be used to weight once again the weighted attention scores a(n,i) or â(n,i) and thus produce the final estimate of the overall delay between the echo and the useful signal picked up by the microphone. This last weighting (416) can be formalized, in the case where the weighted attention scores are the normalized scores â(n,i), by the following equation implemented by a weighting operator of the weighted attention scores. â(n,i) = P(E(n)). â(n,i)

[0117] Unlike standard attention scores, the scores â(n,i) do not sum to 1 and can all be null. This makes it possible, in the absence of echo where P(E(n) «: 1, to indicate to the echo canceller that the loudspeaker signal should not be taken into account when calculating the mask m(n) and thus to avoid degradation of the useful signal.

[0118] In [Fig.6] and [Fig.8] are illustrated the same examples as in [Fig.5] and [Fig.7] but using a modified attention mechanism in accordance with the proposed technique rather than a standard attention mechanism. First of all, we notice that when ¢(.)=0 then the attention score is this time close to zero (608), that is to say that there is nothing to realign (606), resulting in Dr(n) close to zero, where the classic attention module tried at all costs to realign the first signal with the second, even if they have nothing in common (506, 508). Then, we observe in [Fig.8] that the attention score (808) is much more stable over time, essentially conveying the information of delay (around 100 ms) and presence or absence of echo in the microphone, as expected.

[0119] The foregoing description presents an attention score filter which, combined with a standard attention mechanism, forms a modified attention module, as well as an acoustic echo presence detector.

[0120] The proposed attention score filter or the modified attention module integrating such a filter can be integrated into any application requiring signal rephasing.

[0121] The acoustic echo presence detector is a particular example of a detector for a particular application. Generally, the attention score filter can be associated with a module for detecting an expression of a first signal in a second signal decomposed into a plurality of frames.

[0122] In the particular case where the detection module is an acoustic echo presence detector, the first signal is the signal emitted by the loudspeaker, the second signal is the signal picked up by the microphone and the expression of the first signal and the expression of the first signal in the second signal represents the acoustic echo.

[0123] In addition to its application in acoustic echo cancellation, the proposed technique also finds its use in determining the origin of two signals captured separately. For example, the estimated overall delay between the expression of a first signal and that of a second signal can serve as input data for a triangulation mechanism, aiming to locate a spatial and / or temporal origin common to the two signals. This approach is particularly relevant in fields such as the registration of two electrocardiograms obtained via sensors placed on different parts of the body, medical diagnosis by ultrasound, geolocation and structural diagnosis by seismographic techniques.

[0124] Another possible use of the proposed technique concerns signal synchronization. The estimated global delay can be used to temporally coordinate two signals, as in the registration of two videos of the same scene captured by separate cameras. In this case, the dimensions of the convolutional layers of the registration module must be adapted to the two-dimensional data. Moreover, this method finds practical applications in the multimedia field, such as lip synchronization in video conferencing or synchronization of audio tracks in film post-production, thus ensuring increased consistency and quality of the user experience.

[0125] Although the proposed technique has been described in detail with reference to certain preferred embodiments, various modifications and variations may be made without departing from the scope of the invention as defined in the appended claims. For example, the described procedures may be performed in a different order, elements may be added or deleted, and features of different embodiments may be combined as appropriate.

[0126] [Fig.9] represents a device (900), or a system, for estimating the delay between a first signal and a second signal.

[0127] It includes in particular: an output module (902), a receiving module (904), a processor (906) or signal processing unit, and a memory (908) or data storage unit adapted to store instructions of a computer program which, when the program is implemented by the processor, result in implementing the proposed technique.

[0128] Such a device or system may be embedded in a terminal equipped with a loudspeaker configured to emit the first signal and a microphone configured to receive the second signal. In such a terminal, the loudspeaker is then connected to the output module and the microphone is connected to the reception module to allow, for example, the implementation of an echo cancellation mechanism according to an exemplary embodiment.

[0129] The device or system may not include an output module and the receiving module (904) may be configured to receive both the first signal and the second signal. The receiving module may for example be connected to a first sensor configured to receive the first signal and to a second sensor configured to receive the second signal, the first sensor and the second sensor being placed at separate locations. Such a configuration may allow, for example, the implementation of a mechanism for determining an origin of the first signal and the second signal according to an exemplary embodiment. Alternatively, the receiving module can be connected to at least one sensor configured to receive the first signal and the second signal separately, for example at separate times. Such a configuration can allow, for example, the implementation of a method for synchronizing the first signal and the second signal according to an exemplary embodiment.

Claims

Claims

1. Method for estimating a delay between an expression of a first signal and an expression of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the method comprising, for a current frame of the second signal: obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, and obtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally close to the current frame,a weighting (406) of the attention scores with the obtained weights to produce a set of weighted attention scores concentrated around a specific delay corresponding to an estimated overall delay between the expression of the first signal and that of the second signal.,

2. A method according to claim 1, wherein the expressions of the two signals are normalized so that their norm is 1.

3. The method of claim 1, wherein the obtained set of weighted attention scores is normalized to correspond to a set of probabilities.

4. A method of detecting an expression of a first signal in a second signal decomposed into a plurality of frames, the method comprising an estimation of a delay between the expression of the first signal and that of the second signal according to the method of one of claims 1 to 3 and further comprising a determination of a probability of presence of the expression of the first signal in the second signal on the basis of a processing of the sets of weighted attention scores.

5. The method of claim 4, wherein determining the probability of presence comprises implementing one-dimensional convolutional layers acting on a delay axis, each convolutional layer being configured to perform a subsampling of the delay axis.

6. The method of claim 5, wherein combining the sets of weighted attention scores comprises, after the implementation of convolutional layers, an implementation of a pooling layer configured to make the probability of presence of the expression of the first signal within the second independent of the number of frames forming the given temporal context

7. A method of cancelling an echo represented by an expression of a first signal within a second signal, the method comprising an estimation of a delay between the expression of the first signal and an expression of the second signal according to the method of one of claims 1 to 3 or a detection of the expression of the first signal in the second signal according to the method of one of claims 4 to 6.

8. A method of determining an origin of a first signal and a second signal, the first signal and the second signal being captured separately, on the basis of a delay between an expression of the first signal and an expression of the second signal, the method comprising an estimation of the delay according to the method of one of claims 1 to 3 or a detection of the expression of the first signal in the second signal according to the method of one of claims 4 to 6.

9. A method of synchronizing a first signal and a second signal, the method comprising estimating a delay between an expression of the first signal and an expression of the second signal according to the method of one of claims 1 to 3 or detecting the expression of the first signal in the second signal according to the method of one of claims 4 to 6.

10. System for estimating a delay between an expression of a first signal and an expression of a second signal, the first signal and the second signal being temporally decomposed into a plurality of frames, the system comprising the following modules applied to a current frame of the second signal: a module for obtaining a set of unweighted attention scores, where each score represents a probability of correspondence between the current frame considered and a frame of the first signal, a module for obtaining a set of weights, derived from a combination of unweighted attention scores obtained for a given temporal context comprising the current frame and at least one frame of the second signal temporally neighboring the current frame, and a module for weighting (406) the attention scores with the weights obtained by the module for obtaining a set of weights to produce a set of weighted attention scores concentrated around of a specific delay corresponding to an estimated overall delay between the expression of the first signal and that of the second signal.

11. A computer program comprising instructions which, when the program is implemented by a processor, lead to implementing the method according to one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice chip and electronic equipment

    CN109584896A

  • Temporal alignment of signals using attention

    WO2023219751A1