Estimating an optimized mask for processing acquired sound data
By estimating a weighting mask using the sound source's direction of arrival and applying spatial filters, the method enhances distant sound recording, effectively reducing noise and improving speech recognition without neural networks.
Patent Information
- Application Number
- EP2022714494
- Authority / Receiving Office
- EP · EP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-01
- Filing Date
- 2022-03-18
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2042-03-18
AI Technical Summary
Distant sound recording introduces artifacts such as reverberation and ambient noise, degrading the intelligibility of the speaker's voice and complicating communication with speech recognition engines, despite the use of microphone antennas for signal enhancement.
A method that estimates a weighting mask in the time-frequency domain using the direction of arrival of the sound source, applying spatial filtering to enhance the desired signal without relying on neural networks, utilizing techniques like Delay and Sum, MPDR, and MWF filters.
This approach achieves effective noise reduction and signal enhancement with low latency, enabling accurate speech recognition even in adverse conditions, without the high computational costs associated with neural networks.
Smart Images

Figure IMGF0001 
Figure IMGF0002 
Figure IMGF0003
Abstract
Description
Domaine technique
[0001] This description concerns the processing of sound data, particularly in the context of distant sound recording.
[0002] Far-field audio recording occurs, for example, when a speaker is far from a recording device. However, it offers advantages such as real ergonomic comfort for the user to interact hands-free with a service in use: making a phone call, issuing voice commands via a smart speaker (Google Home®, Amazon Echo®, etc.).
[0003] However, this distant sound recording introduces certain artifacts: reverberation and ambient noise appear amplified due to the user's distance. These artifacts degrade the intelligibility of the speaker's voice, and consequently, the functionality of the services. Communication becomes more difficult, whether with a human or a speech recognition engine.
[0004] Hands-free devices (such as smart speakers or teleconferencing "cowboys") are also generally equipped with a microphone antenna that enhances the desired signal by reducing interference. Antenna-based enhancement uses spatial information encoded during multichannel recording and specific to each source to distinguish the signal of interest from other noise sources.
[0005] Numerous antenna processing techniques exist, such as a "Delay and Sum" filter that performs purely spatial filtering using only the arrival direction of the source of interest or other sources, or an "MVDR" (Minimum Variance Distortionless Response) filter that is slightly more effective because it requires knowledge of the spatial distribution of the noise in addition to the arrival direction of the source of interest. Even more powerful filters, such as multichannel Wiener filters, also require the spatial distribution of the source of interest.
[0006] In practice, knowledge of these spatial distributions is derived from a time-frequency map that indicates the points on this map dominated by speech and the points dominated by noise. The estimation of this map, also called a mask, is generally inferred by a pre-trained neural network.
[0007] The following is noted: x ( t , f ) = s ( t, f ) + n ( t, f ) a signal that contains a mixture consisting of both speech and noise in the time-frequency domain, where s ( t , f ) is speech and n ( t , f ) the noise.
[0008] A mask, noted m̂ s ( t , f ) (respectively m̂ n ( t, f) ) , is defined as a real number, usually in the interval [0; 1], such that it is an estimate of the signal of interest ŝ ( t, f ) (respectively noisen ( t, f )) is obtained by simply multiplying this mask by the observations x ( t, f ), either : s ^ t f ≈ m ^ s t f x t f n ^ t f ≈ m ^ n t f x t f
[0009] We are then looking for an estimate of masks m̂ s ( t , f ) And m̂ n ( t , f ), which can lead to the derivation of separation or enhancement filters that are effective. Technique antérieure
[0010] The use of deep neural networks (using an approach implementing "artificial intelligence") has been employed for source separation. A description of such an implementation is presented, for example, in the document [@umbachChallenge], the references for which are given in the appendix below. Architectures such as the simplest "Feed Forward" (FF) type have been investigated and have demonstrated their effectiveness compared to signal processing methods, generally model-based (as described in the reference [@heymannNNmask]). "Recurrent" architectures of the "LSTM" (Long-Short Term Memory, as described in [@laurelineLSTM]) or Bi-LSTM (as described in [@heymannNNmask]) type, which allow for better exploitation of the temporal dependencies of signals, show better performance, in exchange for a very high computational cost.To reduce this computational cost, whether for training or inference, convolutional neural network (CNN) architectures have been successfully proposed ([@amelieUnet], [@janssonUnetSinger]), improving performance and reducing computational cost, with the added possibility of parallelizing calculations. While artificial intelligence approaches for separation generally exploit features in the time-frequency domain, purely temporal architectures have also been used successfully ([@stollerWaveUnet]).
[0011] All these AI-powered enhancement and separation approaches demonstrate real added value for tasks where noise is a problem: transcription, recognition, and detection. However, these architectures share a high cost in terms of memory and computing power. Deep neural network models are composed of dozens of layers and hundreds of thousands, or even millions, of parameters. Furthermore, their training requires large, comprehensive, annotated datasets recorded under realistic conditions to ensure generalization to all usage scenarios.
[0012] A method for processing an audio signal is also known from US patent application 2016 / 0086602A1, which applies spatial filtering to an input signal and a weighting mask to the filtered signal. The mask can be obtained based on spatial selectivity between the signal of interest and the noise of the signal of interest. Résumé
[0013] The present invention improves the situation.
[0014] Claim 1 proposes a method for processing sound data acquired by a plurality of microphones, in which: From the sound data acquired by the plurality of microphones, we determine the arrival direction of a sound from at least one acoustic source of interest, we apply a spatial filtering to the sound data as a function of the arrival direction of the sound, we estimate in the time-frequency domain ratios of a quantity representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other, as a function of the estimated ratios, we develop a weighting mask to apply in the time-frequency domain to the acquired sound data to construct an acoustic signal representing the sound from the source of interest (and thus enhanced relative to ambient noise).
[0015] Here, the "representative quantity" of a signal amplitude refers to the signal's amplitude, but also its energy, power, etc. Thus, the aforementioned ratios can be estimated by dividing the amplitude (or energy, or power, etc.) of the signal represented by the filtered sound data by the amplitude (or energy, or power, etc.) of the signal represented by the acquired (i.e., raw) sound data.
[0016] The weighting mask thus obtained is then representative, at each time-frequency point of the time-frequency domain, of a degree of preponderance of the acoustic source of interest, relative to ambient noise.
[0017] The weighting mask can be estimated to directly construct an acoustic signal representing the sound from the source of interest, enhanced relative to ambient noise, or to calculate second spatial filters which may be more effective at reducing noise more strongly than in the aforementioned case of direct construction.
[0018] In general, it is then possible to obtain a time-frequency mask without using neural networks, with only the following knowledge a priori The direction of arrival of the useful source. This mask then allows the implementation of efficient separation filters such as the MVDR (Minimum Variance Distortionless Response) filter or those from the family of multichannel Wiener filters. The real-time estimation of this mask allows the derivation of low-latency filters. Furthermore, its estimation remains effective even under adverse conditions where the signal of interest is obscured by surrounding noise.
[0019] In implementation, the first spatial filtering mentioned above (applied to the data acquired before estimating the ratios) can be of the "Delay and Sum" type.
[0020] In practice, successive delays can be applied to the signals captured by microphones arranged along an antenna, for example. Since the distances between the microphones, and therefore the phase shifts inherent to these distances between the captured signals, are known, all these signals can be phased and then summed.
[0021] In the case of transforming acquired signals in the ambisonic domain, the signal amplitude represents the phase shifts inherent to the distances between microphones. Here again, it is possible to weight these amplitudes to implement a processing method that can be described as "Delay and Sum".
[0022] In one variation, this initial spatial filtering can be of the MPDR (Minimum Power Distortionless Response) type. It has the advantage of better reducing ambient noise while keeping the useful signal intact, and requires no information other than the direction of arrival. This type of process is described, for example, in the document [@gannotResume], the contents of which are detailed later and whose full reference is given in the appendix.
[0023] Here, however, the MPDR-type spatial filtering, noted w MPDR , can be given in a particular realization by: w MPDR = R x − 1 a s a s H R x − 1 a s , Or a s represents a vector defining the direction of arrival of the sound (or "steering vector"), and R x is a spatial covariance matrix estimated at each time-frequency point ( t, f ) by a relationship of the type: R x t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f x t 1 f 1 x t 1 f 1 H Or : Ω ( t , f) is a neighborhood of the time-frequency point ( t, f ), card is the "cardinal" operator, x (t 1, f 1) is a vector representing the sound data acquired in the time-frequency domain, and x ( t 1, f 1) H< its Hermitian conjugate.
[0024] Furthermore, as previously mentioned, the process may optionally include a subsequent step of refining the weighting mask to denoise its estimation.
[0025] To carry out this subsequent step, the estimation can be denoised by smoothing, for example by applying local averages, defined heuristically.
[0026] Alternatively, this estimate can be denoised by defining a model a priori mask distribution.
[0027] The first approach allows for low complexity, while the second approach, based on a model, achieves better performance, at the cost of increased complexity.
[0028] Thus, in a first embodiment, the weighting mask developed can be further refined by smoothing at each time-frequency point by applying a local statistical operator, calculated on a time-frequency neighborhood of the time-frequency point ( t, f ) considered. This operator can take the form of an average, a Gaussian filter, a median filter, or other.
[0029] In a second embodiment, to carry out the second approach mentioned above, the weighting mask developed can be further refined by smoothing at each time-frequency point, by applying a probabilistic approach comprising: consider the weighting mask as a random variable, define a probabilistic estimator of a model of the random variable, seek an optimum of the probabilistic estimator to improve the weighting mask.
[0030] Typically, the mask can be considered as a uniform random variable in an interval [0,1].
[0031] The probabilistic estimator of the mask M s ( t, f ) can, for example, be representative of a maximum likelihood, based on a plurality of observations of a pair of variables s ^ i x i i = 1 I representing respectively: an acoustic signal ŝ i resulting from the application of the weighting mask to the acquired sound data, and the acquired sound data x i , said observations being chosen from a neighborhood I from the time-frequency point ( t, f ) considered.
[0032] These two implementations are thus intended to refine the mask after its estimation. As mentioned previously, the resulting mask (optionally refined) can be applied directly to the acquired (raw, microphone-captured) data or used to construct a second spatial filter to be applied to this acquired data.
[0033] Thus, in this second case, the construction of the acoustic signal representing the sound from the source of interest and enhanced relative to ambient noise, may involve the application of a second spatial filtering, obtained from the weighting mask.
[0034] This second spatial filtering can be of the MVDR type for "Minimum Variance Distortionless Response", and in this case, at least one spatial covariance matrix is estimated. R n of the ambient noise, the MVDR-type spatial filtering being given by w MVDR = R n − 1 a s a s H R n − 1 a s , with : R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H Or : Ω (t , f ) is a neighborhood of a time-frequency point ( t, f ) , card is the "cardinal" operator, x ( t 1, f 1) is a vector representing the sound data acquired in the time-frequency domain, and x ( t 1, f 1) H< its Hermitian conjugate, and M s ( t 1, f 1) is the expression of the weighting mask in the time-frequency domain.
[0035] Alternatively, the second spatial filtering can be of the MWF type for "Multichannel Wiener Filter", and in this case spatial covariance matrices are estimated R s And R n, respectively of the acoustic signal representing the sound from the source of interest, and of the ambient noise, the MWF type spatial filtering being given by: w MWF = R s + R n − 1 R s e 1 , où e 1 = 1 0 … 0 T , with : R s t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f M s t 1 f 1 x t 1 f 1 x t 1 f 1 H R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H Or : Ω (t , f ) is a neighborhood of a time-frequency point ( t, f ), card is the "cardinal" operator, x ( t 1, f 1) is a vector representing the sound data acquired in the time-frequency domain, and x ( t 1, f 1) H< its Hermitian conjugate, and M s ( t 1, f 1) is the expression of the weighting mask in the time-frequency domain.
[0036] The spatial covariance matrix R n The above represents "ambient noise." This may actually include emissions from sound sources that were not identified as the sound source of interest. Separate processing can be performed for each source whose direction of arrival has been detected (for example, in dynamic mode), and in the processing for a given source, the emissions from other sources are considered part of the noise.
[0037] This implementation illustrates how spatial filtering, such as MWF, can be derived from masking estimates for the most advantageous time-frequency points, where the acoustic source of interest is predominant. It should also be noted that two joint optimizations can be performed, one for the covariance. R s of the acoustic signal involving the desired time-frequency mask M s and the other for covarianceR n ambient noise involving a mask M n related to noise (by selecting time-frequency points in which noise alone is predominant).
[0038] The solution described above thus allows, in general, to estimate in a time-frequency domain an optimal mask in the time-frequency points where the source of interest is predominant, from the sole information of the direction of arrival of the source of interest, without the contribution of a neural network (either to apply the mask directly to the acquired data, or to construct a second spatial filtering to apply to the acquired data).
[0039] The present invention also proposes a computer program (claim 12) comprising instructions for implementing all or part of a method as defined herein when this program is executed by a processor. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.
[0040] The present invention also proposes a device (claim 13) comprising (as illustrated in the figure 3 ) at least one audio data receiving interface (IN) for data acquired by a plurality of microphones (MIC) and a processing circuit (PROC, MEM) configured for: From the sound data acquired by the plurality of microphones, determine an arrival direction of a sound from at least one acoustic source of interest, apply to the sound data a spatial filtering function of the arrival direction of the sound, estimate in the time-frequency domain ratios of a quantity representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand, and as a function of the estimated ratios, develop a weighting mask to be applied in the time-frequency domain to the acquired sound data to construct an acoustic signal representing the sound from the source of interest (and thus enhanced relative to ambient noise).
[0041] Thus, the device may also include an output interface (reference OUT of the figure 3 ) to deliver this acoustic signal. This OUT interface can be connected to a speech recognition module, for example, to correctly interpret user commands despite ambient noise, the delivered acoustic signal having then been processed according to the method described above. Brève description des dessins
[0042] Other features, details, and advantages will become apparent upon reading the detailed description below and analyzing the attached drawings, on which: Fig. 1 [ Fig. 1 ] schematically shows a possible context for implementing the process presented above. Fig. 2 [ Fig. 2 ] illustrates a succession of steps that a process within the meaning of this description may involve, according to a particular embodiment. Fig. 3 [ Fig. 3 ] schematically shows an example of a sound data processing device according to one embodiment. Description des modes de réalisation
[0043] With further reference to the figure 3 Here, the processing circuit of the DIS device presented previously can typically include a MEM memory capable of storing, in particular, the instructions of the aforementioned computer program, as well as a PROC processor capable of cooperating with the MEM memory to execute the computer program.
[0044] Typically, the OUT output interface can feed a voice recognition MOD module of a personal assistant capable of identifying in the aforementioned acoustic signal a voice command from a user UT which, as illustrated in the figure 1 The system can pronounce a voice command captured by a microphone antenna (MIC), even in the presence of ambient noise and / or sound reverberations (REV) generated by the walls and / or partitions of a room, for example, where the user (UT) is located. The processing of the acquired sound data, as described herein and detailed below, nevertheless overcomes such difficulties.
[0045] An example of an overall process as defined in this description is illustrated on the figure 2 The process begins with a first step S1 of acquiring the sound data captured by the microphones. Next, a time-frequency transform is performed on the acquired signals in step S3, after apodization carried out in step S2. The direction of arrival of the sound from the source of interest (DoA) can then be estimated in step S4, specifically by determining the vector a s ( f) of this direction of arrival (or "steering vector"). Then, in step S5, initial spatial filtering is applied to the sound data acquired by the microphones, for example, in the time-frequency domain, and according to the direction of arrival (DoA). This initial spatial filtering can be of the Delay and Sum or MPDR type and is "centered" on the DoA. In the case of an MPDR filter, the acquired data expressed in the time-frequency domain, in addition to the DoA, are used to construct the filter (arrow illustrated with dashed lines for this purpose). Then, in step S6, amplitude (or energy or power) ratios are estimated between the filtered acquired data and the raw acquired data (denoted x ( t, f (in the time-frequency domain). This estimation of the ratios in the time-frequency domain allows us to construct a first, approximate form of the weighting mask, already favoring DoA at step S7 because the aforementioned ratios are high, primarily in the DoA arrival direction. A subsequent, optional step S8 can then be planned, consisting of smoothing this first mask to refine it. Then, at step S9 (also optional), it is possible to generate a second spatial filter from this refined mask. This second filter can then be applied in the time-frequency domain to the acquired audio data in order to generate, at step S10, an acoustic signal substantially free of noise, which can then be properly interpreted by a speech recognition module or other application. Each step of this process is detailed below.
[0046] The following is noted x ( t) an antenna signal composed of N channels, organized as a column vector in step S1: x t = x 0 t ⋮ x N − 1 t
[0047] This vector is called "observation" or "mixture".
[0048] The signals x i , 0 ≤ i < N, These can be signals captured directly by the antenna's microphones, or a combination of these microphone signals as in the case of an antenna collecting signals according to an ambiphonic (also called "ambisonic") format representation.
[0049] In the following, the different quantities (signals, covariance matrices, masks, filters) are expressed in a time-frequency domain, in step S3, as follows: x t f = F x 0 t f ⋮ F x N − 1 t f Or F . is, for example, the short-term Fourier transform of size L : x t f = F x t f = ∑ k = 0 L − 1 x ˜ L , M t − k e − 2 iπkf / L , 0 ≤ f < L
[0050] In the previous relationship, x̃ L,M ( t) is a potentially apodized version at stage S2 by a window w ( k ) and completed with 0s of the variable x ( t ): x ˜ L , M t − k = x t − k w k , si k < M 0 , sinon with M ≤ L and where w ( k ) is a Hann-type or other apodization window.
[0051] Several enhancement filters can be defined depending on the available information. These can then be used for mask deduction in the time-frequency domain.
[0052] For a source s of a given position, we note a s The column vector that points in the direction of this source (the direction of arrival of the sound), a vector called the "steering vector". In the case of a uniform linear antenna made of N sensors, where each sensor is spaced a distance apart from its neighbor d the steering vector of a plane wave at an angle of incidence θ relative to the antenna is defined in step S4 in the frequency domain by: a s f = 1 e − 2 πfd cos θ / c ⋮ e − 2 πf N − 1 d cos θ / c , where c is the speed of sound in air.
[0053] The first channel here corresponds to the last sensor encountered by the sound wave. This steering vector then gives the direction of arrival of the sound or "DOA".
[0054] In the case of a first-order 3D ambisonic antenna, typically in SID / N3D format, the steering vector can also be given by the relation: a s f = 1 cos θ cos ϕ sin θ cos ϕ sin ϕ , where the couple ( θ , ϕ ) corresponds to the azimuth and elevation of the source relative to the antenna.
[0055] Based solely on the knowledge of the direction of arrival of a sound source (or DOA), at step S5 we can define a delay-and-sum (DS) type filter that points in the direction of this source, as follows: w DS = ( a s H< a s -1< a s , Or (.) H< is the transpose-conjugate operator of a matrix or vector.
[0056] A slightly more complex, but also more efficient, filter can be used, such as the MPDR filter (for "Minimum Power Distortionless Response"). This filter requires, in addition to the direction of arrival of the sound emitted by the source, the spatial distribution of the mixture x across its spatial covariance matrix R x : w MPDR = R x − 1 a s a s H R x − 1 a s , where the spatial covariance of the multidimensional signal captured by antenna x is given by the following relation: R x = E xx H
[0057] Details of such an implementation are described in particular in the reference [@gannotResume] specified in the appendix.
[0058] Finally, if we have spatial covariance matrices R s And R n signal of interest s and noise n, we can use a family of much more efficient filters to apply the second spatial filtering mentioned above (described later with reference to step S9 of the figure 2 ). We simply indicate here that, as an example, a spatial filter of the MWF type for "Multichannel Wiener Filter", given by the following equation, can be used as a second filter: w MWF = R s + R n − 1 R s e 1 , où e 1 = 1 0 … 0 T , and involving spatial covariance matrices representing the spatial distribution of acoustic energy emitted by a source of interest R s or by ambient noise R n and propagating through the acoustic environment. In practice, the acoustic properties—reflection, diffraction, scattering—of the materials of the surfaces encountered by sound waves—walls, ceilings, floors, windows, etc.—vary significantly depending on the frequency band considered. Furthermore, this spatial distribution of energy also depends on the frequency band. In addition, in the case of moving sources, this spatial covariance can vary over time.
[0059] One way to estimate the spatial covariance of the mixture x is to perform a local time-frequency integration: R x t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f x t 1 f 1 x t 1 f 1 H Or Ω ( t, f ) is a more or less wide neighborhood around the time-frequency point ( t, f ), And card is the "cardinal" operator.
[0060] From this point, it is already possible to estimate the first filtering w MPDR which can be applied to step S5.
[0061] For matrices R s And R n The situation is different because they are not directly accessible from the observations and must be estimated. In practice, a mask is used. M s ( t, f ) (respectively M n ( t, f )) which allows us to "select" the time-frequency points where the useful source (respectively the noise) is predominant, which then allows us to calculate its covariance matrix by classical integration, by weighting with an appropriate mask of the type: R s t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f M s t 1 f 1 x t 1 f 1 x t 1 f 1 H
[0062] The mask of noise M n ( t, f ) can be derived directly from the useful mask (i.e., associated with the source of interest) M s ( t , f ) by the formula: M n ( t, f) = 1 - M s ( t , fIn this case, the spatial covariance matrix of noise can be calculated in the same way as that of the useful signal, and more specifically in the form: R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M n t 1 f 1 x t 1 f 1 x t 1 f 1 H
[0063] The goal here is to estimate these time-frequency masks M s ( t , f) and M n ( t, f).
[0064] The direction of arrival of the sound (or "DOA", obtained in step S4), originating from the useful source s at time t, is considered known. t, rated doa s ( t This DOA can be estimated by a localization algorithm such as "SRP-phat" ([@diBiaseSRPPhat]), and followed by a tracking algorithm such as a Kalman filter. It can be composed of a single component, as in the case of a linear antenna, or azimuth and elevation components ( θ , ϕ ) in the case of a spherical antenna of the ambisonic type, for example.
[0065] Thus, based solely on the known DOA of the useful source s, we aim in step S7 to estimate these masks. We obtain an enhanced version of the useful signal in the time-frequency domain. This enhanced version is obtained by applying a spatial filter in step S5. w s which points in the direction of the useful source. This filter can be of the Delay and Sum type, or below of the type w MPDR Presented by: w s t = R x − 1 a s a s H R x − 1 a s , si a s t existe 0 , sinon
[0066] From this filter, the signal of interest s is enhanced by applying the filter at step S5: s ^ t f = w s H t f x t f
[0067] This enhanced signal allows us to calculate a preliminary mask M ^ s 0 at stage S7, given by the ratios of stage S6: M ^ s 0 t f = s ^ t f x ref t f γ , Or x ref is a reference channel resulting from the capture, and γ a real positive. γ typically takes integer values (e.g., 1 for amplitude or 2 for energy). It should be noted that when γ → ∞ , the mask tends towards the binary mask indicating the preponderance of the source over the noise.
[0068] For example, for an ambisonic antenna, the first channel, which is the omnidirectional channel, can be used. In the case of a linear antenna, it could be the signal corresponding to any sensor.
[0069] In the ideal case where the signal is perfectly enhanced by the filter w s , And γ = 1, this mask corresponds to the expression: M 0 = s t f s t f + n t f This defines a mask with the desired behavior, namely close to 1 when the signal is predominant, and close to 0 when noise is predominant. In practice, due to the effect of acoustics and measurement imperfections in the source's DOA, the enhanced signal, although already in better condition than the acquired raw signals, may still contain noise and can be further improved by refining the mask estimation (step S8).
[0070] The S8 mask refinement step is described below. Although this step is advantageous, it is not essential and can be carried out optionally, for example if the mask estimated for filtering in step S7 turns out to be noisy beyond a chosen threshold.
[0071] To reduce mask noise, a smoothing function is applied. soft ( .), at step S8. Applying this smoothing function can be reduced to estimating a local average, at each time-frequency point, for example as follows: soft y t f = 1 card Ω 1 t f ∑ t 1 f 1 ∈ Ω 1 t f y t 1 f 1 , Or Ω 1 ( t , f ) defines a neighborhood of the time-frequency point under consideration ( t, f ) .
[0072] Alternatively, one can choose a mean weighted by a Gaussian kernel, for example, or a median operator which is more robust to outliers.
[0073] This smoothing function can be applied either to observations ( ŝ, x ref ), or to the filter M ^ s 0 , as follows: M ^ s 1 t f = soft s t f soft x ref t f M ^ s 1 t f = soft M ^ s 0 t f
[0074] To improve the estimation, we can apply a first saturation step, which ensures that the mask is indeed in the interval [0,1]: M ^ s 2 t f = min M ^ s 1 t f , 1
[0075] Indeed, the previous method sometimes leads to an underestimation of the masks. It can be useful to "correct" the previous estimates by applying a saturation function. sat ( . ) of the type: sat u = u / u th , si u < u th 1 , sinon Or u th is a threshold to be adjusted according to the desired level.
[0076] Another way to estimate the mask from raw observations is, rather than performing averaging operations, to adopt a probabilistic approach, by defining R as a random variable defined by: R = s ^ − M s x , où : ŝ corresponds to the enhanced signal (i.e., filtered by an MPDR or DS enhancement filter), x corresponds to a particular channel of the mixture and Ms corresponds to the mask of the useful source estimated previously: this can be M ^ s 0 or the different variants of M ^ s 1 .
[0077] These variables can be considered as dependent on time and frequency.
[0078] The variable R|M s follows a normal distribution, with a mean of zero and a variance that depends on M s, as follows: R M x ∼ N 0 σ x 2 V R M = V s ^ + M x 2 V x ⇒ σ 2 = σ s 2 + M x 2 σ x 2 Or V (. ) is the variance operator.
[0079] We can also assume a distribution a priori for M s. Since it is a mask, with values between 0 and 1, we assume that the mask follows a uniform distribution in the interval [0,1]: M x ∼ U 0 1
[0080] We can define another distribution that favors the sparsity of the mask, such as an exponential law for example, in a variant.
[0081] From the model imposed for the described variables, the mask can be calculated using probabilistic estimators. Here, we describe the mask estimator M. s ( t , f ) in the sense of maximum likelihood.
[0082] We assume that we have a certain number of observations. Iof the pair of variables s ^ i x i i = 1 I For example, one can select a set of observations by choosing a time-frequency block around the point ( t, f ) where M is estimated s ( t , f) : s ^ i x i = ∪ t 1 ∈ t − δ t , t + δ t , f 1 ∈ f − δ f , f + δ f M s 0 t f x t 1 f 1 x t 1 f 1
[0083] The likelihood function of the mask is written as: p R M s r m s = ∏ i = 1 I 1 2 π σ e − s ^ i − M s x i 2 2 σ 2
[0084] The maximum likelihood estimator is given directly by the expression M s MLE = − b + b 2 − 4 ac 2 a , with : a = ∑ x i 2 − Iσ x 2 b = − 2 ∑ s ^ i x i c = ∑ s ^ i 2 − Iσ s 2 , Or σ s 2 And σ x 2 are the variances of the variables ŝ i And x i .
[0085] Once again, to avoid values outside the interval [0,1], we can apply a saturation operation of the type: M ˜ s MLE = max min M ^ s MLE 1 , 0
[0086] The probabilistic approach is less noisy than the local averaging method. While it is more complex due to the necessary calculation of local statistics, it exhibits lower variance. This allows, for example, for accurate estimation of masks even in the absence of a useful signal.
[0087] The process can continue at step S9 by developing the second spatial filtering from the weighting mask, which in particular gives the matrix M s (as well as the noise matrix) M n = 1 - M s ) to construct a second filter, for example of the MWF type, by estimating the spatial covariance matrices R s And R n specific to the source of interest and the noise, respectively, and given by: R s t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f M s t 1 f 1 x t 1 f 1 x t 1 f 1 H R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H Or : Ω ( t, f ) is a neighborhood of a time-frequency point Ω ( t, f ), card is the "cardinal" operator, x ( t 1, f 1) is a vector representing the sound data acquired in the time-frequency domain, and x ( t 1, f 1) H< its Hermitian conjugate, and M s ( t 1, f 1) is the expression of the weighting mask in the time-frequency domain.
[0088] MWF-type spatial filtering is then given by: w MWF t f = R s t f + R n t f − 1 R s t f e 1 , où e 1 = 1 0 … 0 T .
[0089] It should be noted, alternatively, that if the second filtering method is of the MVDR type, then the second filtering method is given by w MVDR t f = R n − 1 t f a s a s H R n − 1 t f a s with R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H Or Ω ( t, f ) And card are defined as before.
[0090] Once this second spatial filtering has been applied to the acquired data x ( t, f ), we can apply an inverse transform (from time-frequency space to direct space) and obtain an acoustic signal at step S10 x̂ ( t ) representing the sound from the source of interest, enhanced relative to the ambient noise (typically delivered by the OUT output interface of the device shown in the figure 3 ). Application industrielle
[0091] These technical solutions can be applied, in particular, to speech enhancement using complex filters, for example, MWF filters ([@laurelineLSTM], [@amelieUnet]), which ensures good audio quality and a high rate of automatic speech recognition, without the need for a neural network. This approach can be used for keyword or "wake-up word" detection, or even for transcribing a speech signal. Liste des documents cités
[0092] For the record, the following non-patent elements are cited: [@amelieUnet] : Amélie Bosca et al. "Dilated U-net based approach for multichannel speechenhancement from First-Order Ambisonics recordings". In:Computer Speech& Language(2020), pp. 37-51 [@laurelineLSTM] : L. Perotin et al. "Multichannel speech separation with recurrent neuralnetworks from high-order Ambisonics recordings". In:Proc. of ICASSP.ICASSP 2018 - IEEE International Conference on Acoustics, Speech andSignal Processing. 2018, pp. 36-40. [@umbachChallenge] : Reinhold Heab-Umbach et al. "Far-Field Automatic Speech Recognition". arXiv:2009.09395v1. [@heymannNNmask] : J. Heymann, L. Drude, and R. Haeb-Umbach, "Neural network based spectral mask estimation for acoustic beamforming," in Proc. of ICASSP, 2016, pp. 196-200. [@janssonUnetSinger] : A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, "Singing voice separation with deep U-net convolutional networks," in Proc. of Int. Soc. for Music Inf. Retrieval, 2017, pp. 745-751. [@stollerWaveUnet] : D.Stoller, S. Ewert, and S. Dixon, "Wave-U-Net: a multi-scale neural network for end-to-end audio source separation," in Proc. of Int. Soc. for Music Inf. Retrieval, 2018, pp. 334-340. [@gannotResume] : Sharon Gannot et al. "A Consolidated Perspective on Multimicrophone Speech Enhancement and Source Separation". In:IEEE / ACM Transactions on Audio, Speech, and Language Processing25.4 (Apr. 2017), pp. 692-730.issn: 2329-9304.doi:10.1109 / TASLP.2016.2647702. [@diBiaseSRPPhat] : J. Dibiase, H. Silverman, and M. Brandstein, "Robust localization in reverberant rooms," in Microphone Arrays: Signal Processing Techniques and Applications. Springer, 2001, pp. 157-180.
Claims
1. Method for processing sound data acquired by a plurality of microphones (MIC), wherein: - based on the sound data acquired by the plurality of microphones, a direction of arrival of a sound generated by at least one acoustic source of interest is determined; - spatial filtering is applied to the sound data depending on the direction of arrival of the sound; - ratios of a quantity representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand, are estimated in the time-frequency domain; - depending on the estimated ratios, a weighting mask to be applied in the time-frequency domain to the acquired sound data to construct an acoustic signal representing the sound generated by the source of interest is generated.
2. Method according to any of the preceding claims, wherein the spatial filtering is "delay and sum" spatial filtering.
3. Method according to Claim 1, wherein the spatial filtering is applied in the time-frequency domain and is MPDR (Minimum Power Distortionless Response) spatial filtering.
4. Method according to Claim 3, wherein the MPDR spatial filtering, denoted wMPDR, is given by w MPDR t f = R x − 1 t f a s a s H R x − 1 t f a s ,, where as is a vector defining the direction of arrival of the sound, and Rx(t, f) is a matrix of the spatial covariance estimated at each time-frequency point (t,f) by a relationship of the type: R x t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f x t 1 f 1 x t 1 f 1 H where: - Ω(t, f) is a neighbourhood of the time-frequency point (t, f), - card is the "cardinal" operator, - x(t1, f1) is a vector containing the sound data acquired in the time-frequency domain, and x(t1, f1)H its Hermitian conjugate.
5. Method according to any of the preceding claims, wherein the generated weighting mask is further refined by smoothing at each time-frequency point by applying a local statistical operator calculated in a time-frequency neighbourhood of the time-frequency point (t, f) in question.
6. Method according to any of Claims 1 to 4, wherein the generated weighting mask is further refined by smoothing at each time-frequency point, and wherein a probabilistic approach is applied, the probabilistic approach comprising: - considering the weighting mask to be a random variable, - defining a probabilistic estimator of a model of the random variable, - seeking an optimum of the probabilistic estimator to improve the weighting mask.
7. Method according to Claim 6, wherein the mask is considered to be a uniform random variable in an interval [0, 1].
8. Method according to either of Claims 6 and 7, wherein the probabilistic estimator of the mask Ms(t, f) is representative of a maximum likelihood, over a plurality of observations of a pair of variables s ^ i x i i = 1 I , namely respectively: - an acoustic signal ŝi resulting from application of the weighting mask to the acquired sound data, and - the acquired sound data xi, said observations being selected in a neighbourhood of the time-frequency point (t,f) in question.
9. Method according to the preceding claims, wherein construction of the acoustic signal representing the sound generated by the source of interest and enhanced with respect to ambient noise comprises application of second spatial filtering based on the generated weighting mask.
10. Method according to Claim 9, wherein the second spatial filtering is MVDR (Minimum Variance Distortionless Response) spatial filtering, and at least one spatial covariance matrix Rn(t,f) of the ambient noise is estimated, the MVDR spatial filtering being given by w MVDR t f = R n − 1 t f a s a s H R n − 1 t f a s , with: R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H where: - Q(t,f) is a neighbourhood of a time-frequency point (t,f), - card is the "cardinal" operator, - x(t1, f1) is a vector containing the sound data acquired in the time-frequency domain, and x(t1, f1)H its Hermitian conjugate, and - Ms (t1, f1) is the expression of the weighting mask in the time-frequency domain.
11. Method according to Claim 9, wherein the second spatial filtering is MWF (Multichannel Wiener Filter) spatial filtering, and spatial covariance matrices Rs and Rn are estimated for the acoustic signal representing the sound generated by the source of interest, and the ambient noise, respectively, the MWF spatial filtering being given by wMWF(t, f) = (Rs(t, f) + Rn(t, f))-1Rs(t, f)e1, où e1 = [1 0 . . . 0]T, with : R s t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f M s t 1 f 1 x t 1 f 1 x t 1 f 1 H R n t f = 1 card Ω t f ∑ t 1 f 1 ∈ Ω t f 1 − M s t 1 f 1 x t 1 f 1 x t 1 f 1 H where: - Q(t,f) is a neighbourhood of a time-frequency point (t,f), - card is the "cardinal" operator, - x(t1, f1) is a vector containing the sound data acquired in the time-frequency domain, and x(t1, f1)H its Hermitian conjugate, and - Ms(t1,f1) is the expression of the weighting mask in the time-frequency domain.
12. Computer program comprising instructions for implementing the method according to any of the preceding claims when the program is executed by a processor.
13. Device comprising at least one interface (IN) for receiving sound data acquired by a plurality of microphones (MIC) and a processing circuit (PROC, MEM) configured to: - based on the sound data acquired by the plurality of microphones, determine a direction of arrival of a sound generated by at least one acoustic source of interest; - apply spatial filtering to the sound data depending on the direction of arrival of the sound; - estimate in the time-frequency domain ratios of a quantity representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand; and - depending on the estimated ratios, generate a weighting mask to be applied in the time-frequency domain to the acquired sound data to construct an acoustic signal representing the sound generated by the source of interest.
Citation Information
Patent Citations
Sound signal processing method, and sound signal processing apparatus and vehicle equipped with the apparatus
US20160086602A1
Speech enhancement method and system, computer equipment and storage medium
CN110503972A
Multichannel noise cancellation using deep neural network masking
US10522167B1
Enhancement of audio from remote audio sources
US20210082450A1