Obtaining an impulse response of a room

EP4634624A1Pending Publication Date: 2025-10-22ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023818504
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-12-07
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing methods for obtaining an impulse response of a room are limited by the complexity of inferring acoustic properties from sound signals, especially in unfavorable conditions, and are hindered by noise and the computational complexity of real-time processing, with conventional microphone devices often restricted to first-order ambisonic formats that suffer from noise amplification and spatial aliasing.

Method used

A method involving a time-frequency transform to express the generalized velocity vector in the frequency domain, followed by an inverse transform to model it in the time domain using an autoregressive moving average (ARMA) filter, allowing for a more robust extraction of the impulse response that characterizes the acoustic space, even with lower ambisonic orders, and optimizing the filter to model the series of peaks representing direct and reflected sound paths.

Benefits of technology

This approach enables a more general and robust characterization of acoustic spaces, providing a reduced room impulse response that is causal and relative, effectively capturing reflection delays and amplitudes, thus improving the accuracy of sound source localization and separation in various applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

It is proposed to process sound signals acquired by an array of microphones and generated by a sound source, with a view to acoustically characterizing a space (ESP) comprising the array (MIC) and the source and bounded by a wall (PAR). A time-frequency transform is applied to the acquired signals, and a generalized velocity vector is expressed in the frequency domain on the basis of the acquired signals. In particular, the generalized velocity vector is expressed in the time domain v(t) in the form of a succession of peaks comprising at least one peak associated with reflection from the wall and with a time abscissa dependent on the delay TAU1, and the expression in the time domain of the generalized velocity vector is modelled by an autoregressive moving average ARMA defined by an autoregressive filter AR and a moving average MA. Thus, provision is made to process the acquired sound signals by applying the autoregressive filter AR, to obtain an impulse response characterizing the space (ESP) and obtained from the moving average MA.
Need to check novelty before this filing date? Find Prior Art

Description

Description Title: Obtaining an impulse response from a room Technical field

[0001] This description relates to the field of sound data processing. It relates more particularly to obtaining an impulse response of a room (partitioned space), from an impulse response called a “generalized relative impulse response”. Prior art

[0002] Knowledge of the acoustic and geometric properties of an environment can help obtain or improve the obtaining of relevant results in the processing of audio signals for a multitude of use cases. It can be advantageous to simultaneously perform audio processing including both the localization and the separation of sound sources in an environment, particularly in unfavorable conditions (for example in the presence of obstacles preventing sound propagation in a straight line). The needs for such processing are numerous, particularly in applications of spatial encoding, augmented reality, robot navigation, room characterization, and others.

[0003] When sound is used to estimate the acoustic environment, it is generally necessary to exploit the characteristics of multi-microphones that encode spatial information. A particularly well-suited representation of a 3D sound field is the Higher Order Ambisonics (HOA) format, hereinafter called "ambisonics", based on the decomposition of the acoustic pressure on a sphere into spherical harmonics. Ambisonic channels coincide with each other but differ in their directivity, i.e., in their sensitivity to excitations coming from different spatial directions. They can be recorded by specific devices (most often spherical microphone arrays called "SMAs") or artificially created.In a given environment, and for the given source and SMA microphone device position, each HOA channel admits a particular room impulse response (denoted "RIR"). These RIR responses provide information on the environment where the sound propagates, in particular in the first part of the responses (i.e. in the "first echoes").

[0004] Even if the spatial fingerprint is embedded in the recorded audio, retrieving this information is not straightforward. On the one hand, RIRs are related to the (unknown) source signal, and on the other hand, the recording may be contaminated by noise. For this reason, all inference methods base their analysis on a set of pre-recorded or estimated RIRs (not necessarily in HOA format). While such an analysis can be challenging in itself (e.g., due to the echo “labeling” problem of attributing each signal peak to a reflection from a partition wall), this is a very strong assumption that limits applications only to use cases where RIRs are available. To circumvent this problem, an inference approach blind can be based on the analysis of so-called "phase-aligned" spatial correlation matrices. However, the computational complexity of this approach seems prohibitive for real-time processing.

[0005] Alternatively, one could consider relative fingerprints, i.e., a relative transfer function (denoted “ReTF”, in the frequency domain) or a relative impulse response (denoted “ReIR”, in the time domain) to infer the properties of the environment. ReTF and ReIR model the relationship between individual channels and a given reference signal, which is usually chosen to be one of the channels. Theoretically, these representations are source-independent, but the price to pay in this method is that some information is inevitably lost (in particular, the propagation time and absolute attenuation of a signal propagating directly from the source to the microphone). Typically, ReIR responses are not causal and their analysis is much more complex than that of RIRs.

[0006] In the work corresponding to document WO-2022 / 106765, it has nevertheless been demonstrated that the use of the reference signal which is a linear combination of all channels (i.e. a reference beamform) is advantageous when extracting information from the ReIR of ambisonic signals. More particularly, if beamforming (hereinafter beamforming) sufficiently attenuates the acoustic reflections compared to direct propagation, the corresponding ReIR (called “Generalized Velocity Vector”, and noted “GTVV” hereinafter) admits an informative and compact expression in the time domain. Under these conditions, the generalized ReIR is causal and relatively sparse, and therefore allows an estimation, based on the peak of the direction of arrival of the sound (noted DoA), of the directions of the acoustic reflections and their associated delays.

[0007] The GTVV vector (i.e., ReIR in the beamforming ambisonic domain) is more robust to adverse acoustic conditions than the “standard” ReIR, for which the reference signal is usually the zero-order omnidirectional ambisonic channel. However, it is limited by the performance of the applied beamforming. For example, if a maximum directivity, signal-independent beamforming is used, its directivity is a quadratic function of the given HOA order. However, conventional SMA microphone devices generally do not provide sufficiently high-order ambisonic formats: most often, they are only able to record first-order ambisonic signals (FOA). This is especially true for simple devices, e.g., handheld devices, supporting FOAs only.Furthermore, the frequency support of higher-order channels gradually decreases with increasing HOA order, as noise amplification at low frequencies and spatial aliasing at high frequencies begin to manifest.

[0008] However, the favorable theoretical properties of the GTVV vector tend to diminish at low ambisonic orders, due to the inability of beamforming to effectively suppress reflections. The problem is further compounded by increasing the distance between the microphone and the source, as more reflections fall into the main beamforming lobe and into the at the same time the preponderance of direct sound decreases with respect to reflections. In practice, we can observe that the GTVV imprint is no longer causal, and that the estimated directions are less precise.

[0009] Moreover, even when the GTVV representation remains valid, extracting directions and delays by peak identification / selection is not necessarily straightforward. An easy-to-use GTVV vector can be considered as the multichannel RIR response (without delay, centered) implied by a causal filter. The consequence is that the same reflection is infinitely repeated as an echo, at time instances corresponding to integer multiples of its relative delay, with its sign alternating and amplitude decreasing. Thus, these series can interfere with each other, altering the information that can be deduced from them, or even masking the presence of reflections of lower amplitudes for example. Abstract

[0010] This description improves the situation.

[0011] To this end, it proposes a method for processing sound signals acquired by at least one array of microphones and originating from at least one sound source, to acoustically characterize a space comprising the array and the source and delimited by at least one wall, in which: - A time-frequency transform is applied to the acquired signals, - From the acquired signals, a generalized velocity vector V(f) is expressed in the frequency domain, complex with a real part and an imaginary part, the velocity vector characterizing a composition between: * a first acoustic path, direct between the source and the array of microphones, represented by a first vector U0, and * at least one second acoustic path originating from a reflection on the wall and represented by a second vector U1, the second path having, at the array of microphones, a delay TAU1, relative to the direct path, - An inverse transform is applied, from frequencies to time,to the generalized velocity vector to express it in the time domain v(t) in the form of a succession of peaks comprising at least one peak linked to the reflection on said wall and to a time abscissa which is a function of the delay TAU1.,

[0012] In particular in this method, the expression in the time domain of the generalized velocity vector is modeled by an autoregressive moving average ARMA defined by an autoregressive filter AR and a moving average MA, and the method then comprises processing of the sound signals acquired by application of the autoregressive filter AR, to obtain an impulse response characterizing said space and resulting from the moving average MA.

[0013] Thanks to this arrangement, the information stored in the representation of the generalized velocity vector, expressed in the time domain (and hereinafter noted as "GTVV"), can be extracted in a more robust manner because it is more general for any acoustic situation, in order to obtain an impulse response characterizing a space with at least one wall (a space such as a room and thus corresponding to an impulse response of the RIR type for “Room Impulse Response”). More particularly, as described later in the exemplary embodiments, this impulse response can be described as “reduced” (and noted “RdRIR” for “Reduced Room Impulse Response”) because the temporal expression of the generalized velocity vector, from which this impulse response is deduced, only presents reflection delays relative to the reception delay at the microphone of the direct acoustic path from the source (and not delays in absolute terms). Similarly, the amplitudes of the reflections are relative to the amplitude of reception at the microphone of the direct sound (not reflected by a wall).Nevertheless, such an impulse response, even relative, already makes it possible to effectively characterize the acoustic space considered, simply by treating the temporal expression of the generalized velocity vector as an ARMA model.

[0014] Thus, this reduced impulse response RdRIR comes from the ARMA model, and is distinguished in this from the relative impulse response ReIR, introduced previously, which is obtained directly from the expression of the generalized velocity vector.

[0015] In one embodiment, the acquired signals are applied to ambisonic channels, and the autoregressive AR filter is common to all channels.

[0016] Such an implementation in ambisonic representation has the advantage of not requiring too high an ambisonic order (first orders or "FOA" for "First Order Ambisonic" being sufficient to obtain a satisfactory impulse response).

[0017] In an embodiment where the aforementioned space is delimited by a plurality of walls, the expression in the time domain of the generalized velocity vector comprises a series of peaks comprising a peak linked to the direct path (or "DoA" for "Direction of Arrival") followed by peaks each linked to at least one reflection on a wall n. The method then comprises: - optimizing the autoregressive filter to model said series of peaks in the form of a multivariate autoregressive moving average.

[0018] Thus, the temporal representation of the generalized velocity vector presents itself well to modeling by a multivariate ARMA average.

[0019] In such an embodiment in particular, the method may comprise: - from the expression of the generalized velocity vector in the time domain v(t) in the form of said series of peaks, optimizing the autoregressive filter by exploiting a causality property of an impulse response.

[0020] Indeed, the generalized velocity vector can be expressed in the time domain in the form: where and are causal filters representing the moving average part MA respectively and the regressive part AR of the ARMA model, and are linked by , for a beamforming w to be received by the microphone array and according to a direction of arrival of the sound from the sound source, and where: , ^^ ^^ ∈ ]0,1[ and denote the parameters of an nth plane wave reflected by a wall n of space, ^^ ^^ being a directional encoding vector of propagation of the nth wave, ^^ ^^ being a relative attenuation of the nth wave and ^^ ^^ being a delay of the nth wave with respect to said direct path, ^^0 being a propagation vector specific to the direct path, being a beamforming response to the nth wavefront, with .

[0021] In such an implementation, the autoregressive part can then be estimated by minimizing: is a channel of the generalized velocity vector, represented by a multivariate ARMA model, the autoregressive part being common to all channels of the generalized velocity vector.

[0022] In such an implementation, the estimate by minimizing ^^ comes down to solving a linear prediction system, which is advantageously overdetermined.

[0023] With the notations presented above, the impulse response can be given by the moving average , such that:

[0025] Since the time representation of the generalized velocity vector can have both positive and negative amplitudes (as illustrated as an example in Figure 2), an amplitude sign correction can be applied to the moving average to obtain the usual expression of the impulse response: in positive form.

[0026] Furthermore, in one embodiment, said impulse response is chosen to have a finite duration. This duration may be chosen in particular to avoid taking into account a diffuse reverberation field (typically high-order multiple reflections which would appear on the far right of Figure 2) and thus only process early reflections on the wall(s) of the space considered.

[0027] In such an embodiment, this property according to which the aforementioned impulse response is of finite duration can be exploited to set a maximum limit on the length of filters for the autoregressive part AR and for the moving average part MA.

[0028] In a further embodiment, the moving average portion MA may be centered on a delay corresponding to a time of reception at the microphone of the sound from the source.

[0029]

[0030] According to another aspect, there is provided a computer program comprising instructions for implementing the above method, when these instructions are executed by a processing circuit. According to another aspect, there is provided a non-transitory recording medium, readable by a computer, on which such a program is recorded.

[0031] According to another aspect, there is also provided a device comprising a processing circuit comprising an interface for receiving data from sound signals acquired by a microphone network, and configured to implement the above method. Brief description of the drawings

[0032] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which:

[0033] Figure 1 illustrates an example of a sequence of steps in a process of the above type,

[0034] Figure 2 illustrates an example of a time representation of the generalized velocity vector,

[0035] Figure 3 shows real examples of generalized velocity vectors (center) under different conditions, the corresponding ARMA representations (right), and the real impulse responses (left), with ARMA models being more faithful to the real impulse responses,

[0036] Figure 4, Figure 5 and Figure 6 illustrate the performance of results obtained by implementing the above method compared to other treatments (or absence of treatment), respectively for different durations of sound reverberation cycles,

[0037] Figure 7 illustrates an error evaluation on the sound arrival direction (DoA) only, under experimental conditions similar to those of Figures 4 to 6, showing a clearer performance of the implementation of the above method on reflections in particular,

[0038] Figure 8 schematically illustrates a device for implementing the method. Description of the embodiments

[0039] We first refer to figure 1 illustrating steps of a method of the above type according to an exemplary embodiment.

[0040] We first describe the main principles of the steps in Figure 1.

[0041] Ambisonic signals of any order are for example recorded by a SMA microphone device (or are generated otherwise, by simulation or otherwise). These multichannel signals are then used for the estimation of the generalized velocity vector GTVV. This estimation is often conveniently carried out by calculating the inverse Fourier transform of the corresponding relative transfer function ReTF in the frequency domain, as described notably in document WO-2022 / 106765. Thus, ambisonic signals are generally transformed into a time-frequency representation (e.g. by STFT, for “Short-Time-Fourier-Transform”) beforehand, and a robust estimator is used to obtain the ReTF.

[0042] The temporal shape of the GTVV vector is presented for illustrative purposes in Figure 2 and shows: - a peak at zero delay and linked to the direct acoustic path, associated with the main DoA of the sound (for "Direction of Arrival"), and - peaks linked to higher delays and linked to reflections on partitions.

[0043] The temporal footprint of the GTVV vector can be considered as the realization of a multivariate autoregressive moving average or “ARMA” process, where the autoregressive AR filter (in the denominator) is common to all channels.

[0044] Therefore, once the GTVV vector is obtained, the calculations are carried out by estimating the parameters of the corresponding ARMA model. The AR filter is first estimated from the time series given by the GTVV vector as illustrated as an example in Figure 2, taken from the aforementioned document WO-2022 / 106765, by exploiting the fact that the RIR responses are causal (as is the case for the AR and MA filters of the ARMA model).

[0045] One can further set the maximum limit on the length of the AR and MA filters, since the first part of the RIRs is assumed to have a finite duration; in practice, such an AR filter can be efficiently computed by estimating a linear prediction model applied to the appropriate part of the GTVV fingerprint.

[0046] Once the AR filter is available, the MA filters can be estimated by simply convolving the GTVV vector by the AR filter (an efficient estimate, such as Prony), or by least-squares estimation (such as Shanks estimation); here too, more advanced inference procedures can be considered, for example by applying some structure between the corresponding inputs of the MA filters.

[0047] Ideally, MA filters should approximate normalized RIR responses, whose main peak is centered at zero delay of the representation (invariant to absolute gain and delay, due to the information loss in ReIR, as mentioned earlier). Due to its similarity to RIRs, such a sequence of MA filters is called RdRIR for "Reduced Room Impulse Response". In reality, RIR responses are considered to have positive amplitude and are continuous functions over time. When represented by a discrete (multichannel) time series, RIR responses are pre-filtered by an anti-aliasing filter, which often presents an impulse response containing both positive and negative amplitudes.Since the same filter is applied to all HOA channels, it is possible to observe the sign of the estimated zero-order RdRIR response and (if negative) reverse the sign of all channels for a given time sample.

[0048] After correcting the signs of the RdRIR representation, we can proceed to infer the acoustic wavefronts. A non-limiting way to do this is to perform peak selection on the series of amplitude peaks of the RdRIR response at different times.

[0049] The application of ARMA modeling to the GTVV (generalized ReIR) vector in the ambisonic domain is described in more detail below.

[0050] The principles presented can be adapted to standard (non-ambisonic) ReIR responses, for example, for the purpose of estimating the time difference of arrival (or TdoA), considering a pair of microphones recording the same source signal.

[0051] The following mathematical description covers the definition of the GTVV vector, the wavefront inference when the latter is theoretically valid (i.e., when the convergence condition explained below is satisfied), as well as the derivation of the ARMA-based “GTVV preconditioning” method presented above.

[0052] We note below the vector of concatenated spherical harmonic expansion coefficients (denoted "SH") (corresponding to the "HOA channels") up to order L, at frequency f. The recorded signals are assumed to be due to a far-field sound source at azimuth , elevation and distance from the SMA microphone array, in an indoor environment (typically a partitioned room). Given a broadband beamforming w (or "beamforming" hereinafter) directed (approximately) towards the DoA, the generalized velocity vector in the frequency domain (GFVV) is defined as follows, as described in particular in WO-2022 / 106765:

[0054] where, ^^ ^^ ∈ ]0,1[ and denote the parameters of the nth plane wave reflected by a partition of the room, with:

[0055] , the SH expansion vector in the direction

[0056] ^^ ^^ , its relative attenuation and

[0057] ^^^^ , its delay (compared to the direct propagation component).

[0058]

[0059] Then, is the SH vector of the plane wave in the DoA direction given by while is the response of the beamformers to the nth wavefront (with ).

[0060]

[0061] The approximation is due to the simplifying assumptions built into the right-hand side of the above equation: the plane wave decay has been given in terms of dominant acoustic reflections, and the beamforming and relative attenuations are assumed to be independent of frequency.

[0062]

[0063] The inverse Fourier transform per channel of the vector GFVV gives its temporal counterpart GTVV:

[0064]

[0065] In practice, the processing is done in the STFT (Short-Time-Fourier-Transform) domain, and the GTVV time duration is dictated by the chosen window. The window length is centered relative to the GTVV, at , that is,

[0066] Under the condition of convergence of the (geometric) Taylor series, the GTVV admits an expression of the form: accumulates ''cross terms'' (which relate to the mutual interference between different wavefronts).

[0067] The above expression "Equation1" allows us to immediately estimate the wavefront of the direct sound by evaluating , while the remainder involves the summation of the infinite series corresponding to the reflected wavefronts.

[0068] But since , each infinite series has an amplitude which decreases with the time position of the peak:

[0070] When beamforming is very selective, its response is . If this is not the case, we can improve the estimate of by “debiasing” the observed vector

[0071]

[0072] Given an estimate of , and a collection of SH vectors (corresponding to a set of directions ), and knowing that is strictly positive, we can recover by finding an element that maximizes the correlation with in the Equatio, which Equation 3

[0073] Alternatively, one can resort to nonlinear optimization and solve Equation 2 in parametric form, where becomes the function of the direction variables .

[0074]

[0075] The convergence condition: implies that beamforming significantly attenuates reflections, which depends on the type of beamforming applied, but also of course on the acoustic environment and the HOA order.

[0076] For computational reasons, it is convenient to use simple beamformings, such as the maximum directivity beamforming given by (in N3D ambisonic encoding, knowing that it is sufficient to weight the signals acquired by the ambisonic microphone (with several piezoelectric capsules to collect several sound signals) to move from one type of encoding to another).

[0077] However, due to the width of its main lobe, this beamforming is too permissive at low ambisonic orders (e.g. the FOA(s)), and therefore, the expression in Equation 1 may no longer be valid.

[0078]

[0079] However, the GTVV vector can still be written in the form: where and are both causal filters, linked by

[0080]

[0081] This expression reveals a particular structure (each GTVV channel can be seen as a realization of the multivariate ARMA model, whose autoregressive part is common to all channels).

[0082] The MA part, or the series of the RdRIR (Reduced Room Impulse Response) response, thus admits an expression of the type:

[0083] Since , we can estimate by minimizing under the constraint

[0084]

[0085] This is advantageously an overdetermined problem: the length of the filter is , while the number of data points is (the non-causal part of the GTVV vector representation). With increasing order of HOA, the estimation should become more accurate, as more data become available for regression.

[0086]

[0087] This cost function can be extended to incorporate weights, as well as the last part of the wavefront series, which is assumed to be a low-magnitude noise-like signal:

[0088] As the two filters are related by the linear expression , by imposing the condition , the filter support is also implicitly shortened by has .

[0089] In principle, it would be possible to integrate more structure into (or ), by further modifying the original cost function. One such example might be to use norms favoring the sparseness of a group to model the support However, solving such an optimization problem usually requires additional computational resources. Therefore, a least-squares minimization is proposed here as an example.

[0090] Taking the partial derivative of ^^ with respect to a filter element AR and noted , and setting the result to zero, we obtain:

[0092] The two autocorrelation functions defined above can be efficiently calculated using a fast Fourier transform. Their overall (weighted) sum can be denoted:

[0094] Since the estimation of the remaining coefficients amounts to a classic linear prediction problem: which can be solved by various methods.

[0095] For example, in order to use fast Toeplitz-based solvers, it is possible to slightly modify the original cost function and instead minimize a substitute function of the type:

[0100] Once has been calculated, we can retrieve the non-zero segment of (the RdRIR) by evaluating.

[0101]

[0102] Such an implementation is very computationally efficient. However, one can choose to apply a more elaborate approach such as estimating the RdRIR in the least-squares sense (the so-called "Shank method"), or even performing an alternating optimization to improve both the AR and the RdRIR (known as the Steiglitz-McBride algorithm). These approaches require the estimation of the inverse AR filter, which is usually approximated by an optimal FIR filter in the least-squares sense.

[0103]

[0104] The feature representation is given in matrix form where represents the sequence of GTVV vectors from Equation 1 or the estimated RdRIR sequence from Equation 3, for each .

[0105]

[0106] An example of such sequences, for a recording from an SMA device collecting FOAs from a speech source, and for a real multichannel RIR response (shifted so that its main peak is placed at ) is given in Error! Source of the cross-reference not found.. It appears that the proposed RdRIR then more closely approximates the RIR structure than the fingerprint of the GTVV vector.

[0107]

[0108] The DoA is evaluated from the vector corresponding to zero delay in the matrix, while the remaining directions are obtained by selecting the amplitude peaks of its column vectors. The index of the chosen peak reveals the relative delay of the given direction with respect to the direct sound path.

[0109] Then, an angular error on the directions associated with the ten largest peaks of the corresponding sequence can be quantified.

[0110]

[0111] Three approaches are considered below: - none (no post-processing), - unbiased (bias correction using Equation 3), and - with arma correction (RdRIR) in the sense of the process presented above, and this for different ambisonic orders (or "order") , , or .

[0112]

[0113] Specifically, the evaluations are conducted for the order HOA , , and , with SNR equal to 0dB, 10dB, 20dB and “Inf” dB (i.e., practically noise-free). Each result is the median estimate of 10 repetitions of the given simulation setup (i.e., for the given reverberation time and additive white Gaussian noise level). The experiments simulate a rectangular room of size 5 x 4 x 3 m 3 , with the microphone array and voice source positioned randomly, but their distance between them being between 0.5 and 6 m.

[0114]

[0115] The experiment implementation for three reverberation cycles (RT60=200ms, RT60=400ms and RT60=600ms) is presented in Error! Reference source not found., Error! Reference source not found. and 6 respectively. The results clearly present that RdRIR provides the most accurate estimates, with the performance of all approaches increasing with the HOA order, and worsening with increasing reverberation time and noise level. It It is striking, however, that RdRIR often outperforms the remaining approaches, even when its HOA order is lower than that of the other two approaches.

[0116]

[0117] In Error! Reference source not found. an evaluation of the DoA error only, under similar experimental conditions, for RT60 = 400ms, is presented. Although the RdRIR estimate again has the lowest angular error, for all SNR levels, the difference here is less significant. This suggests that the main improvement of ARMA post-processing lies in the better prediction of the wavefronts that are reflected in particular.

[0118]

[0119] Figure 8 illustrates an example of a device for implementing the above method, and typically comprising: - an interface INT for receiving signals from a microphone MIC, for example ambisonic (with several piezoelectric capsules for example), the microphone MIC being arranged in an ESP space comprising at least one PAR wall, - a processor PROC connected to the interface INT to process the signals received, for example in ambisonic representation, express the generalized velocity vector in time as a function of these signals, and deduce the ARMA model therefrom to deliver an impulse response RdRIR of the ESP space, - a memory MEM storing instruction data of a computer program within the meaning of the present description, and accessible by the processor PROC to read this data and execute the above method.

[0120]

[0121] Obtaining the impulse response of the ESP space makes it possible in particular to quantify the acoustic and geometric properties of this space (for example to simultaneously obtain the localization and separation of sound sources in the ESP space, or others). Knowledge of the acoustic and geometric properties of such an ESP environment can make it possible to obtain or improve the obtaining of relevant results in the processing of audio signals for various applications of spatial encoding, augmented reality, robot navigation, room characterization, and others. As demonstrated above, the use of the ARMA model to obtain this room impulse response is simple to implement (in particular for the low ambisonic order required) and gives satisfactory results as illustrated in Figures 4 to 7.

Claims

Claims

1. 1. Method for processing sound signals acquired by at least one array of microphones and originating from at least one sound source, to acoustically characterize a space comprising the array and the source and delimited by at least one wall, in which: - A time-frequency transform is applied to the acquired signals, - From the acquired signals, a generalized velocity vector V(f) is expressed in the frequency domain, complex with a real part and an imaginary part, the velocity vector characterizing a composition between: * a first acoustic path, direct between the source and the array of microphones, represented by a first vector U0, and * at least one second acoustic path originating from a reflection on the wall and represented by a second vector U1, the second path having, at the array of microphones, a delay TAU1, relative to the direct path, - An inverse transform is applied, from frequencies to time,to the generalized velocity vector to express it in the time domain v(t) in the form of a succession of peaks comprising at least one peak linked to the reflection on said wall and to a time abscissa which is a function of the delay TAU1, in which the expression in the time domain of the generalized velocity vector is modeled by an autoregressive moving average ARMA defined by an autoregressive filter AR and a moving average MA, The method comprising processing the acquired sound signals by applying the autoregressive filter AR, to obtain an impulse response characterizing said space and resulting from the moving average MA.

2. 2. Method according to claim 1, in which the acquired signals are applied to ambisonic channels, the autoregressive filter AR being common to all the channels.

3. 3. Method according to one of the preceding claims, in which, for a space delimited by a plurality of walls,the expression in the time domain of the generalized velocity vector comprises a series of peaks comprising a peak linked to the direct path (DoA) followed by peaks each linked to at least one reflection on a wall n, The method comprising: - optimizing the autoregressive filter to model said series of peaks in the form of a multivariate autoregressive moving average.

4. 4. The method of claim 3, comprising: - from the expression of the generalized velocity vector in the time domain v(t) in the form of said series of peaks, optimizing the autoregressive filter by exploiting a causality property of an impulse response.

5. 5. The method of claim 4, wherein the generalized velocity vector is expressed in the time domain in the form:, and the regressive part AR of the ARMA model, and are linked by , for a beamforming w to be received by the microphone array and according to a direction of arrival of the sound from the sound source, and where: , ^^ ^^ ∈ ]0,1[ and denote the parameters of an nth plane wave reflected by a wall n of space, ^^ ^^ being a directional encoding vector of propagation of the nth wave, ^^ ^^ being a relative attenuation of the nth wave and ^^ ^^ being a delay of the nth wave with respect to said direct path, ^^0 being a propagation vector specific to the direct path, being a beamforming response to the nth wavefront, with .

6. 6. The method of claim 5, wherein the autoregressive part is estimated by minimizing: is a channel of the generalized velocity vector, represented by a multivariate ARMA model, the autoregressive part being common to all channels of the generalized velocity vector.

7. 7. The method of claim 6, wherein the estimation of by minimizing ^^ amounts to solving a linear, overdetermined prediction system.

8. 8. Method according to one of claims 5 to 7, in which the impulse response is given by the moving average such that:

9. 9. Method according to claim 8, in which an amplitude sign correction is applied to the moving average to obtain the impulse response in positive form.

10. 10. Method according to one of the preceding claims, wherein said impulse response is of finite duration.

11. 11. Method according to claim 10, wherein a maximum limit of filter length is set for the autoregressive part AR and for the moving average part MA.

12. 12. Method according to one of the preceding claims, in which the moving average part MA is centered on a delay corresponding to a time of reception at the microphone of the sound from the source.

13. 13. Computer program comprising instructions for implementing the method according to one of the preceding claims, when said instructions are executed by a processing circuit.

14. 14. Device comprising a processing circuit comprising an interface for receiving data from sound signals acquired by a network of microphones, and configured to implement the method according to one of claims 1 to 12.