Improved acoustic source location method

The method enhances acoustic source localization by using a generalized velocity vector to account for direct and reflected sounds, improving accuracy and stability in direction and distance estimation.

JP7771184B2Active Publication Date: 2025-11-17オランジュ
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023530282
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-19
Filing Date
2021-10-15
Publication Date
2025-11-17
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

Existing acoustic source localization methods, particularly those using small microphone systems, struggle with accurate direction and distance estimation due to interference from direct and reflected sounds, especially when placed near walls or reflective surfaces, leading to biased and incomplete estimates.

Method used

A method for processing acoustic signals using a generalized velocity vector that accounts for both direct and indirect paths, and determines the direction and distance of the sound source by applying a time-frequency transformation, followed by an iterative process to refine the direction and distance estimates.

Benefits of technology

The method provides more precise and stable direction and distance estimation by incorporating the interference of direct and reflected sounds, reducing bias and improving accuracy in acoustic source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771184000077
    Figure 0007771184000077
  • Figure 0007771184000078
    Figure 0007771184000078
  • Figure 0007771184000079
    Figure 0007771184000079
Patent Text Reader

Abstract

The present invention relates to the processing of acoustic signals acquired by at least one microphone, for example of the Ambisonic type, in order to localize at least one sound source in a space with at least one wall. A time-frequency transformation is applied to the acquired signals, and based on the acquired signals a general complex velocity vector V(f) with real and imaginary parts is represented in the frequency domain, this vector having components in its denominator other than the omnidirectional component W(f). In particular, this vector is a first acoustic path, represented as a first vector U0, which is the straight line path between the source and the microphone; characterizing the combination of at least a second sound path caused by reflections at the wall and represented as a second vector U1; The second path has a first delay TAU1 at the microphone relative to the straight path. As a function of the delay TAU1, the first vector U0, and the second vector U1, at least one parameter of the direction of the straight line path (DoA), the distance d0 from the source to the microphone, and the distance z0 from the source to the wall is determined.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to acoustic source positioning methods, and in particular to positioning methods for estimating the direction of sound, or "DoA" (direction of arrival), using small microphone systems (e.g., microphones capable of receiving sound, in the "ambiophonic" or "ambisonic" terms used below). [Background technology]

[0002] An example application is beamforming, which involves spatial separation of sound sources, in particular to improve speech recognition (e.g., virtual assistants with voice interaction). Such processing can also be used for spatial audio coding (pre-analysis of an audio scene in order to encode the main signals separately), i.e., to spatially edit immersive audio content, possibly audiovisually (for art, electronic music, film, or other purposes). This also makes it possible to track speech in teleconferences and to detect audio events (with or without associated video).

[0003] In the state of the art for Ambisonic (or equivalent) coding, most techniques are based on spatial components generated by frequency analysis (usually time-frequency representations generated by processing with a short-time Fourier transform, or "STFT"), or representations of narrowband time signals generated by filter banks).

[0004] The first-order Ambisonic signal is collected in vector form according to Equation 1 in the Appendix. For convenience, the encoding convention of Equation 1 is presented here, but is not intended to be limiting, as conversion to other conventions can be implemented as appropriate. Thus, the field equivalent to a single plane wave carrying the emitted signal s1(t) coming from the direction described by the unit vector U1 (and therefore the source direction DoA) can be described by Equation 2 (Appendix).

[0005] In practice, the signal is analyzed frame by frame in the frequency domain to obtain Equation 3 (Appendix), which in the case of a single wave takes the form of Equation 4, and by extension for N waves takes the form of Equation 5.

[0006] Certain methods are based on the analysis of the velocity vector V(f) or the intensity vector I(f), the former being an alternative to the latter normalized by the power of the omnidirectional reference component, as expressed in Equations 6 and 7.

[0007] Methods utilizing complex frequency samples essentially base their position estimation on the information contained in the real part of such vectors (related to the active intensity and wave propagation characteristics directly related to the gradient of the phase field).

[0008] The imaginary part (the reactive part associated with the energy gradient) is considered to be characteristic of stationary acoustic phenomena.

[0009] It can be seen that for a single plane wave the velocity vector is V = U1.

[0010] A known method (known as "DirAC") operates either on time samples filtered into subbands, which are real and also have an intensity vector, or on complex frequency samples, where only the real part of the intensity vector is used to specify the direction of arrival (or, more precisely, its opposite direction). Furthermore, the calculation of a so-called "diffuseness" coefficient, related to the ratio between the norm of the vector and the energy of the sound field, makes it possible to determine whether the information available at a given frequency is instead characteristic of a directional component (where the direction of the vector determines the position) or of "ambience" (a mixture of diffuse reverberant sound and / or undifferentiated secondary sound sources).

[0011] Another method, hereafter referred to as "VVM," is based on statistics of the velocity vector and the angular direction of the real part, weighted by a coefficient related to the ratio of the real and imaginary parts to their norms. A spherical cartography (2D histogram, e.g., equirectangular) is established by collecting values ​​for all frequency samples and for a certain number of time frames. Thus, the estimate is essentially based on maximum probability and subject to a certain latency.

[0012] Another class of methods, sometimes referred to as "covariance" methods, is sometimes presented as a first extension and involves calculating the covariance matrix of the spatial components (sometimes called the power spectral density matrix, or "PSD") in the frequency subbands. Again, the imaginary part may be completely ignored. Note that the first row (or column) of this matrix is ​​equivalent to the magnitude vector if the spatial components are Ambisonic. Many of these techniques involve "subspace" methods and algorithms that can be expensive, especially when operating on many frequency subbands or when using high spatial resolution.

[0013] These "vector-based" or "matrix-based" methods seek to obtain "directional" components associated with locatable acoustic source components, i.e., paths, on the one hand, and with environmental components, on the other hand.

[0014] Among the observed limitations of such methods, they are hindered by the interference of direct sound (which indicates the direction of the acoustic source) with reflected sound, even in the case of a single simultaneous acoustic source. In the presence of certain room effects, for example, accurate estimates are often not obtained and / or the estimates are often too biased. When the object containing the acoustic localization and capture device (e.g., an Ambisonic microphone) is placed, for example, on a table or near a wall (and / or this is the acoustic source), such reflective surfaces are likely to induce an overall angular bias.

[0015] Indeed, localization is generally biased by the total interference of direct sound with reflected sounds associated with the same acoustic source. When based on velocity vectors, it is mainly the real part of the velocity vector that is considered, while the imaginary part is usually ignored (or at least not used to its fullest extent). Acoustic reflections are considered nuisance and are not incorporated into the estimation problem. Thus, without considering the specific interference structures induced, a neglected and unmodeled component remains.

[0016] Thus, in the above types of applications, acoustic positioning is generally estimated in angular terms only. Furthermore, there are unlikely to be any effective methods that offer distance estimation based on a single capture point (which seems characteristic of simultaneous, or more generally, "miniature" microphone systems, i.e., small in size compared to the distance from the acoustic source, typically housed within a volume of about 10 cm for Ambisonic microphones).

[0017] However, in certain applications, additional information about the distance from the source is required in addition to the direction (and therefore the stereo localization in XYZ). These include, for example: - Virtual navigation in a stereoscopically captured real environment (since the proper correction of the source angle and intensity depends on the relative XYZ movement of the object and the microphone), - Source location to identify the speaker (especially for connected speakers or other devices); - Surveillance and alarm systems in residential or industrial environments; others.

[0018] A particularly useful approach, presented in document FR1911723, uses the sound velocity vector to determine the location of the barrier, in particular by obtaining the sound's direction of arrival, its delay (and therefore the distance to the source) and the delay of the reflected sound. Such an implementation makes it possible to model the interference of a direct wave with at least one indirect wave (due to reflections) and to use a representation of the model for the entire velocity vector (its imaginary and real parts).

[0019] However, although this technology is already in use, it remains to be further improved.

[0020] The present invention improves this situation. [Prior art documents] [Patent documents]

[0021] [Patent Document 1] FR1911723 [Non-patent literature]

[0022] [Non-Patent Document 1] Jerome Daniel (2000) Summary of the Invention [Means for solving the problem]

[0023] A method for processing an acoustic signal acquired by at least one microphone is proposed, To locate at least one sound source in a space having at least one wall, - a time-frequency transformation is applied to the acquired signal, Based on the acquired signals, a generalized velocity vector V'(f) is estimated from the equation of the velocity vector V(f), expressed in the frequency domain, in which the reference component D(f) other than the omnidirectional component W(f) appears in the denominator, the equation being a complex equation with real and imaginary parts, and the generalized velocity V'(f) is a first acoustic path, represented as a first vector U0, which is the straight line path between the source and the microphone; characterizing the combination of at least a second sound path caused by reflections at the wall and represented as a second vector U1; The second path has a first delay TAU1 at the microphone, relative to the straight path; as a function of the delay TAU1, the first vector U0, and the second vector U1, *Direction of Arrow (DoA) and *The distance d0 from the source to the microphone, and a distance z0 from the source to the wall, wherein at least one parameter is determined.

[0024] The generalized velocity vector V'(f) mentioned above is thus constructed from the velocity vector V(f), which is generally expressed as a function of the omnidirectional component of the denominator. The generalized velocity vector V'(f) replaces the "conventional" velocity vector V(f) in the sense of the above-mentioned document FR1911723, to obtain a "reference" component other than the omnidirectional component in the denominator. The reference component may actually be more "separable" toward the sound arrival direction. Furthermore, in the exemplary embodiment shown below with reference to Figures 6A to 6D and 7, the sound arrival direction from which the reference component can be calculated is obtained as a first approximation, for example, by using the conventional velocity vector V(f) during the first iteration of an iterative procedure that gradually converges toward the accurate DoA.

[0025] It has been observed within the scope of the present invention that the determination of the above-mentioned parameters DoA, d0, z0 becomes more precise and / or accurate in particular by using such a general velocity vector V'(f) with more relevant reference components instead of the velocity vector V(f). In particular, the method is more stable, especially in situations where acoustically strong reflections come from barriers placed near the microphone or near the active sound source.

[0026] The parameters DoA, d0, and z0 mentioned above are cited above. Note that in the exemplary embodiment, the general velocity vector formula allows determining, in particular, the delay TAU1 mentioned above.

[0027] In one embodiment, the method includes multiple iterations, as described above, using at least in part a generalized velocity vector V'(f) whose denominator includes a reference component D(f) determined based on an approximation of the direction of approach (DoA) obtained in a previous iteration. Under almost all circumstances, these iterations converge to a more accurate DoA.

[0028] Such a method then includes a first iteration in which a "conventional" velocity vector V(f) is used instead of the general velocity vector V'(f). As described in document FR1911723, the velocity vector V(f) is expressed in the frequency domain, with the omnidirectional component W(f) appearing in the denominator. After the first iteration, at least a first approximation of the direction of approach (DoA) can be determined.

[0029] Thus, at least in the second iteration following the first, a general velocity vector V'(f) is used, estimated from the equation for the velocity vector V(f) in which the omnidirectional component W(f) in the denominator is replaced by the reference component D(f), which is more spatially resolved than the omnidirectional component W(f).

[0030] For example, in one embodiment, the reference component D(f) is more highly resolved in a direction corresponding to the above-mentioned first approximation of the direction of the straight path (DoA).

[0031] The iterative process may be repeated until convergence is reached according to a predetermined criterion, which may in particular be a causality criterion to identify with reasonable certainty at least the first reflections off obstacles (i.e., the aforementioned "barriers") in the sound propagation environment between the microphone and the source.

[0032] In a specific embodiment, each iteration comprises: an inverse frequency-to-time transformation is also applied to the expression for the generalized velocity vector V'(f) in order to obtain, in the time domain, a series of peaks, each associated with the reflection of the sound at at least one wall, in addition to the peaks associated with the arrival of the sound along the straight path (DoA); - if a signal appears in a series of peaks whose time coordinates are less than those of a straight line and whose amplitude exceeds a selection threshold, a new iteration is performed (possibly adaptive), The causality criterion is met if the amplitude of the signal is below a threshold. Obtaining such a series of peaks may typically lead to the form shown in equation B4=39b in the appendix, which is further described below with reference to Figure 2, where it should be understood that it applies to the general velocity vector V'(f).

[0033] The iterative process of the above method can be performed, for example, by - in a first case, the amplitude of said signal is less than a selected threshold; - in the second case, when repeated iterations of the process do not lead to a significant decrease in the amplitude of this signal; It may be terminated.

[0034] In one exemplary embodiment, following the second case, the following steps are performed, the acquired signal being obtained in the form of successive frames of samples, each step comprising: - for each frame, a score for the presence of a sound onset in that frame is estimated (e.g., an equation such as Equation 53 in the Appendix); - frames with a score higher than the threshold are selected for processing the acoustic signal acquired within said frame.

[0035] Indeed, if convergence towards a DoA solution is not easy due to the proximity of the barriers from which the first direct reflections originate, it is preferable to obtain direct reflections at these barriers relative to the onset of the sound (when the sound emission begins).

[0036] Considering the respective formulas of the “traditional” and “general” velocity vectors, in an embodiment where the acquired signal is obtained by an Ambisonic microphone, the “traditional” velocity vector V(f) can be represented in the frequency domain by first-order Ambisonic components of the following kind of form: V(f)=1 / W(f)[X(f),Y(f),Z(f)] T , W(f) is the omnidirectional component, In the frequency domain, the generalized velocity vector V'(f), represented by the first order Ambisonic components, can be presented in the following kind of format: V(f)=1 / D(f)[X(f),Y(f),Z(f)] T , D(f) is a reference component other than the above-mentioned omnidirectional components.

[0037] The order considered here is first order, which allows the components of the velocity vector to be expressed in a three-dimensional coordinate system, but other implementations are possible, especially higher Ambisonic orders.

[0038] In one embodiment, an estimate of the direction of a straight path equivalent to the first vector U0 can be determined from the average value over a set of frequencies of the real part of the generalized velocity vector V'(f) expressed in the frequency domain (following the form of Equation 24, which naturally applies here for the generalized velocity vector V'(f)).

[0039] Thus, the expression for the velocity vector in the frequency domain can provide an estimate of the vector U0.

[0040] However, in a further improved embodiment, - an inverse frequency-to-time transform is applied to the general velocity vector to represent it in the time domain V'(t), - After the time of the straight path, at least one Maximum is found as a function of time in the formula for the general velocity vector V'(t)max, From there, a first delay TAU1 is deduced, which corresponds to the time of the maximum value V'(t)max.

[0041] Furthermore, in this embodiment, a second vector U1 is estimated as a function of the values ​​of the normalized velocity vector V' recorded at time indices t=0, TAU1 and 2×TAU1, so as to define a vector V1 as follows: V1=V'(TAU1)-((V'(TAU1)V'(2.TAU1)) / ||V'(TAU1)|| 2 )V'(0) And the vector U1 is U1=V1 / ||V1||.

[0042] and, The angles PHI0 and PHI1 of the first vector U0 and the second vector U1, respectively, can be determined for the wall as follows: PHI0=arcsin(U0.nR) and PHI1=arcsin(U1.nR), where nR is the unit vector, normal to the wall; The distance d0 between the source and the microphone, as a function of the first delay TAU1, is given by a relation of the following kind: It can be determined by d0 = (TAU1xC) / ((cosPHI0 / cosPHI1)-1), where C is the speed of sound.

[0043] Furthermore, the distance z0 from the source to the wall can be expressed by the following type of relationship: z0=d0(sinPHI0-sinPHI1) / 2 It can be determined by

[0044] In this way, all parameters for the positioning of the source can be determined (for example, as in Figure 1), here for the case where there is a single wall, but the model can be generalized to the case where there are several walls.

[0045] Thus, in embodiments where the space comprises multiple walls, - an inverse frequency-to-time transformation is applied to the general velocity vector in order to represent it in the time domain V'(t) in the form of a series of peaks (a form corresponding to the first approach for equation 39b in the appendix); - identifying a peak in the series of peaks associated with a reflection at one wall of the plurality of walls, each identified peak having a time coordinate that is a function of a first delay TAUn of a sound path caused by the reflection at the corresponding wall n relative to a straight line path; at least one parameter as a function of a first vector U0 representing the sound path caused by the reflections at the wall n and each first delay TAUn of each second vector Un, *Direction of Arrow (DoA) and *The distance d0 from the source to the microphone, * At least one distance zn from the source to the wall n.

[0046] As shown in Figure 5B, which is an example applied to "traditional" velocity vectors but is also applicable to general velocity vectors, after inverse transformation (from frequency to time), the velocity vector equations (traditional and general) show a series of peaks, which are also shown in Figure 2 for illustration purposes, where maximum values ​​are reached for multiple values ​​of the above-mentioned delays (TAU1, 2TAU1, etc., TAU2, 2TAU2, etc.) between the straight path and the path caused by at least one reflection of sound at the wall, and combinations of these delays (TAU1+TAU2, 2TAU1+TAU2, TAU1+2TAU2, etc.).

[0047] These peaks can then be used to identify, in particular, a peak relating to at least one reflection at wall n, whereby there are multiple time coordinates (x1, x2, x3, etc.) for the delay TAUn associated with this wall n.

[0048] Since the combination of various delays may make it difficult to identify a single delay (TAU1, TAU2, TAU3, etc.) and the corresponding presence of a wall, a first portion of the peaks with the smallest positive time coordinates may be preselected to identify peaks associated with wall reflections in this portion (avoiding the various delay combinations TAU1+TAU2, 2TAU1+TAU2, TAU1+2TAU2, etc. that may appear after the first delay), although it is assumed in such an implementation that the causality criteria described above are met (alternatively, a "second" peak may also be obtained by combining a delay with a negative multiplier, which in combination with a "positive" delay ultimately results in a small positive time coordinate).

[0049] In this way, a peak associated with a reflection at a wall n may have a time coordinate that is a multiple of the delay TUn associated with this wall n. In the ideal case, the first portion of peaks with the smallest positive time coordinate can be preselected to identify each peak in that portion associated with a single reflection at a wall.

[0050] In embodiments where the signal acquired by the microphone is in the form of a series of samples, it is more generally possible to apply a weighting window with exponential variation that decreases over time to these samples (see Figure 5A and described further below).

[0051] Furthermore, the window can be placed at the beginning of the sound onset (or even before the sound onset), thereby avoiding difficulties with multiple reflections.

[0052] By applying such a weighting window, it is possible to obtain a less biased first estimate of parameters U0, d0, etc. by utilizing the velocity vector expression in the time domain, especially when a "traditional" velocity vector is involved, e.g., in the context of the first iteration of the method. Indeed, in certain situations where the cumulative intensity of reflected sound is greater than the direct sound, the estimates of the above-mentioned parameters may be biased. These situations can be detected when peaks are observed at negative time coordinates in the time representation of the velocity vector (top curve in FIG. 5B). Applying a weighting window of the above kind allows these peaks to return to positive coordinates, as shown in the bottom curve in FIG. 5B, resulting in a less biased estimate.

[0053] However, this implementation is optional if, in this type of situation, a nearly unbiased estimate of the parameters U0, d0, etc. is already possible by using a general velocity vector instead of a "traditional" velocity vector. However, such an operation could also be performed, for example, in the first iteration of the method using a "traditional" velocity vector, or in the second case mentioned above where the iteration is non-convergent.

[0054] In one embodiment, furthermore, weightings q(f) may be applied iteratively to the velocity vectors (general or conventional) in the frequency domain, each associated with a frequency band f, according to the following type of equation: q(f)=exp(-|Im(V(f)).m| / (||Im(V(f))||) where Im(V(f)) is the imaginary part of the velocity vector (conventional or common, here simply referred to as "V(f)"), and m is the unit vector perpendicular to the plane defined by vector U0, which is normal to the wall (typically the Z axis in Figure 1, which is described in more detail below).

[0055] According to such an embodiment, the frequency bands that are most effective for determining the above-mentioned parameters are selected.

[0056] The invention also relates to an acoustic signal processing apparatus comprising a processing circuit adapted to implement the above method.

[0057] For illustrative purposes, FIG. 4 shows a schematic representation of such a processing circuit, which includes:

[0058] an input interface IN for receiving a signal SIG acquired by a microphone (which may, for example, comprise several piezoelectric plates for generating the signal in an Ambisonic manner);

[0059] a processor PROC which processes the signals in cooperation with the working memory MEM and in particular develops the equation of the general velocity vector in order to derive the desired parameters d0, U0 etc., the values ​​of each of which can be provided by an output interface OUT.

[0060] Such a device may take the form of a module for locating a sound source in a stereoscopic environment, connected to a microphone (acoustic antenna or other type), or conversely, in augmented reality, an engine for generating a sound based on a given position of the source in a virtual space (including one or more walls).

[0061] The invention also relates to a computer program comprising instructions for implementing the above method when executed by a processor of a processing circuit.

[0062] For example, Figures 3A, 3B and 7 show example flow charts that may be considered as algorithms for such programs.

[0063] In another aspect, a non-transitory computer readable medium having such a program stored thereon is provided.

[0064] In the detailed description therein, general velocity vectors and "conventional" general vectors are designated with the same V-notation (V(f), V(t)), without distinction being made by the same term "velocity vector", especially in the formulae given in the appendix. Specifically, when referring to general velocity vectors, the velocity is explicitly designated in that term and is written as V' (V'(f), V'(t)). The first part of the description up to Fig. 5B confirms the principles underlying the format used in document FR1911723 and contained here in the formulae given in the appendix.

[0065] More generally, other features, details, and advantages will become apparent from reading the following detailed description and examining the accompanying drawings, which are set forth below. [Brief explanation of the drawings]

[0066] [Figure 1] For illustrative purposes, FIG. 1 shows various parameters involved in sound source localization according to one embodiment. [Figure 2] FIG. 2 shows a series of different peaks, presented for illustrative purposes as a time representation of a velocity vector after an inverse frequency-to-time transform ("IDFT"). [Figure 3A] Figure 3A shows the first few steps in the algorithmic process to determine the relevant parameters U0, d0, etc. [Figure 3B] FIG. 3B shows a continuation of the processing steps of FIG. 3A. [Figure 4] FIG. 4 shows a schematic diagram of an apparatus within the scope of the present invention according to one embodiment. [Figure 5A] FIG. 5A illustrates a weighting window for samples of an acquired signal that decays exponentially over time, according to one embodiment. [Figure 5B] In Figure 5B we compare the post-IDFT time representation of the velocity vector: - without pre-processing of the samples by a weighting window (top curve), - with windowing (bottom curve). [Figure 6A]FIG. 6A shows a summary of relevant peaks in the time representation of the generalized velocity vector V'(t) as the iterative process of the method described below with reference to FIG. 7 is repeated. [Figure 6B] FIG. 6B shows a summary of the relevant peaks in the time representation of the generalized velocity vector V'(t) as the iterative process of the method described below with reference to FIG. 7 is repeated. [Figure 6C] FIG. 6C shows a summary of the relevant peaks in the time representation of the generalized velocity vector V'(t) as the iterative process of the method described below with reference to FIG. 7 is repeated. [Figure 6D] FIG. 6D shows, highly diagrammatically and for illustrative purposes, the form of the reference component D(f) that appears in the denominator of the equation for the general velocity vector V'(f) over several iterations of the method: [Figure 7] FIG. 7 shows diagrammatically the steps of an iterative method within the scope of the invention, according to the embodiment shown here by way of example. DETAILED DESCRIPTION OF THE INVENTION

[0067] The velocity vectors can be calculated by known methods, however certain specific parameter settings may be recommended to improve the final results obtained.

[0068] Typically, the frequency spectrum B(f) of an Ambisonic signal is first obtained by a short-time Fourier transform (i.e., STFT) for a series of time frames b(t), typically overlapping (e.g., overlap / add), where the order of the Ambisonic components may be m=1 for four components (although, without loss of generality, the calculations can also be adapted to higher orders).

[0069] Then, for each time frame, the velocity vector is calculated as follows: - for a conventional velocity vector, as a ratio to the omnidirectional component W(f) (Equation 6 in the Appendix), or - The general velocity vector is calculated as a ratio to the reference component D(f) by replacing W(f) with D(f) in Equation 6 in the Appendix.

[0070] Embodiments can be envisaged that introduce temporal smoothing or integration by weighted sum as described below.

[0071] If, at such ratios (X(f) / W(f), Y(f) / W(f), Z(f) / W(f); X(f) / D(f), Y(f) / D(f), Z(f) / D(f)), the spectral content of the audio signal actually enhances a substantial amount of useful frequencies (e.g., over a wide frequency band), the characteristics of the original signal are substantially removed, in favor of emphasizing the characteristics of the audio channel.

[0072] In the above-mentioned applications, we can consider the situation of an acoustic source of stable characteristics (at least over several consecutive frames, in position and in radiation) emitting a signal s(t) in a stable acoustic environment (reflective walls and objects, possibly diffractive, etc.), thus accounting for what is usually referred to as the "room effect," even if it is not in an actual "room." These signals are captured by an Ambisonic microphone. The Ambisonic signal b(t) results from the so-called "acoustic channel effect," which is a combination of spatial encodings of various versions of the signal s(t) along direct and indirect paths. This results in the convolution of the signal with a spatial impulse response h(t), where each channel (i.e., dimension) is associated with an Ambisonic component, expressed as Equation 8 in the Appendix.

[0073] This impulse response is called SRIR, or "Spatial Room Impulse Response," and is generally represented as a series of temporal peaks: - the first peak located at time t=TAU0 (time of propagation) corresponding to the direct sound, - a second peak at time t=TAU1 corresponding to the first reflection, And so on.

[0074] Thus, at these peaks, the direction of arrival of these wavefronts can be calculated as the vector u given in Equation 9-1 to a first approximation: n In reality, the spatial impulse response is unknown data, but here we explain how we can arrive at some of its features indirectly through the velocity vector calculated based on the Ambisonic signal b(t).

[0075] To emphasize this, we first explain the relationship between the impulse response h(t), the transmitted signal s(t), and the Ambisonic signal b(t) (Equation 9-2) over a selected observation time. To be precise, this equation assumes that there is no measurement noise and that there are no other acoustic sources whose signals are captured, directly or indirectly, over that time. Thus, all direct and indirect signals from the source are captured over this time.

[0076] It has been shown that performing a Fourier transform over this time span, the resulting velocity vector has a unique signature of the spatial impulse response. This so-called LT transform (because it is "longer term" than the STFT) converts b(t), s(t), and h(t) into B(f), S(f), and H(f) according to Equation 10. This temporal subdivision may correspond to a time window extending over several consecutive signal frames.

[0077] The velocity vector equation calculated in Equation 11 is then derived from the convolution equation calculated in the frequency domain. If there is non-zero energy (in fact, detectable energy) for each frequency f over time, then Equation 11 will be characteristic of the acoustic channel (i.e., room effects) and not of the transmitted signal.

[0078] In practice, as mentioned above, a common approach is to perform a frame-by-frame time-frequency analysis, in which each short-time Fourier transform is applied to a time window in which it is not possible in principle to be sure that the observed signal originates entirely and exclusively from the convolution product of Equation 9. This means that, strictly speaking, the velocity vector cannot be described in a form that only characterizes the acoustic channel (as in the right-hand side of Equation 11). However, here, an approximation is made as best as possible within the context of this description (Equation 20, introduced below), and the advantages of the short-time analysis are exploited as follows:

[0079] In a subsequent step, a series of energy peaks characterizing the linear path of the signal emitted by the source and received by the microphone is obtained to capture the first reflections from one or more walls, insofar as these reflections are identifiable, and one can then focus on what characterizes the beginning of the spatial impulse response, i.e., first on the first time peak from which the direction of the direct sound is inferred, and possibly on subsequent time peaks characterizing the first reflections.

[0080] To do this, the effect of the interference of the direct sound with at least one reflected sound is investigated in the complex velocity vector equation to estimate relevant parameters for locating the sound source.

[0081] For the onset of the impulse response shown in Equation 12, we introduce a simple model of a straight path (n=0) combined with N reflections (n=1,...,N). In this equation, g n , TAU n and u n are respectively the attenuation, delay and direction of arrival of the wave with index n (n reflections) before reaching the microphone system. In what follows, for the sake of brevity, but without imposing any restrictions on generality, we consider the delay and attenuation relative to the direct sound, which will result in setting the terms in equation 13 to n=0.

[0082] The corresponding frequency equation is given in Equation 14 for the special case of gamma 0 = 1 for direct sound. n It follows that for any n greater than 0, is a function of frequency f.

[0083] The frequency representation of the Ambisonic field is given by Equation 16, if the latter half is ignored.

[0084] The short-term velocity vector is then expressed as Equation 17, or, normalized with a non-zero epsilon term to (approximately) avoid infinite values ​​when the denominator components are (approximately) zero, as Equation 18. In Equation 17 or Equation 18, the component W (characteristic of the conventional velocity vector) can be replaced by the reference component D to represent the generalized velocity vector. In fact, in the general case, D replaces W in the denominator, and the expression for the conventional velocity vector V corresponds to the special case where D = W. However, for simplicity of explanation here, the notation related to the special case of D = W is presented in the first equation of the appendix for the conventional velocity vector, but can be easily substituted for the generalized velocity vector by noting the replacement of W by D.

[0085] Short-time analysis allows observing the frequency fingerprints (hereafter referred to as "FDVV") characteristic of the wavefront submixes in the spatial impulse response over time and according to the dynamic evolution of the original signal. For a given observation, the characteristic submix (smx is a "submix") is modeled in the time and frequency domain by Equation 19:

[0086] In the method described below, the frequency fingerprint FDVV is calculated as a submix H smxIn fact, for the conventional velocity vector, the expression in Eq. 20 is given here, and H_W is expressed as the following matrix D Θ It can be applied to general velocity vectors by replacing it with the matrix vector H by

[0087] In particular, at the signal onset, the implicit model h smx (t) is the start of the spatial impulse response h early (t), at least in terms of relative wavefront direction and delay. The implicit relative gain parameter g n is affected by the time window and the dynamic characteristics of the signal and does not necessarily correspond to that of the impulse response. Here, we mainly consider situations where the observation is characteristic, focusing mainly on the direct wave (which provides the DoA) and one or several early reflections.

[0088] In particular, for illustrative purposes, an example of a conventional process for estimating velocity vectors in the frequency domain while considering only a single reflection is presented below; the case of a general velocity vector will be described later. Here, we consider the case of a single interference (effectively between a direct sound and the first reflection) to highlight the specific spatial frequency structure and demonstrate how the desired parameters can be determined by considering not only the real part of the velocity vector but also the imaginary part. Indeed, an Ambisonic field is described according to Equation 21, which infers the velocity vector according to Equation 22. From this equation, the real and imaginary parts travel in parallel parts of the three-dimensional space (affine and linear, respectively) as the frequency travels through the relevant acoustic spectrum shown in Equation 23. The affine part (real part) lies on a line containing the unit vectors U0 and U1 representing the direct and indirect waves, respectively, and these two parts are perpendicular to the midplane of these two vectors (thus, the imaginary part of the vector remains unchanged, since it lies on the linear part). Furthermore, statistical calculations assume a uniform distribution of the phase difference between the waves (and thus a representative sweep of frequency), and the average of the real part of the velocity vector is equal to the vector U0 as expressed in Equation 24. The maximum probability is the average of U0 and U1 weighted by the wave amplitudes as shown in Equation 25. Therefore, maximum probability-based DoA detection is hindered by a regular angular bias, giving a direction intermediate between the direct sound and its direction. Equation 23 shows that this spatial sweep is made by periodically equating the frequency to the inverse of the delay TAU1 between the two waves. Therefore, if such a spatial-frequency structure is observable, the directions U0 and U1, as well as the delay TAU1 from the observation, can be extracted. Furthermore, other methods for estimating these parameters in the time domain are presented below (discussed in conjunction with Figure 2).

[0089] By predicting the orientation of the reflecting surface relative to the microphone's reference frame, the estimation of U0, U1, and TAU1 allows us to infer the absolute distance d of the source relative to the microphone, and possibly the altitude of both. Indeed, as shown in Figure 1, if the distance from source S0 to microphone M is d0 and d1 is its mirror image S1 with respect to reflecting surface R, then surface R is perpendicular to the plane formed by vectors U0 and U1. These three points (M, S0, S1) lie in the same plane perpendicular to surface R. The parameters to be determined to define the orientation (i.e., tilt) of the reflecting surface remain specified. In the case of reflections from floors or ceilings (detected this way because U1 points at the floor or ceiling), we can use the assumption that it is horizontal and parallel to the XY plane of the Ambisonic microphone's coordinate system. The distances d0 and d1 are then related by relation 26, which in turn directly gives the distance from microphone M to axis (S0, S1), and PHI0 and PHI1 are the elevation angles of vectors U0 and U1, respectively.

[0090] Also, an estimate of the delay TAU1 of the reflected sound relative to the direct sound is obtained, which allows for another distance-distance relationship 27 to be used, since their difference indicates the delay of the acoustic path, with coefficient c being the speed of sound.

[0091] By expressing d1 as a function of d0, only this last quantity is unknown and can be estimated according to equation 28. We obtain the distance from the source to the reflecting surface, i.e., its height, or altitude above the ground, z0, according to equation 29, and that of the microphone according to equation 30.

[0092] The various parameters U0, U1, PHI0, PHI1, d1, d0, etc. are shown in Figure 1 for the example of sound reflected from a floor. It is obvious that similar parameters can be inferred for sound reflected from a ceiling. Similarly, similar parameters can be inferred for sound reflected from other reflecting surfaces R whose pose relative to the microphone coordinate system is known, and whose pose is determined by the normal n R(unit vector perpendicular to the surface R). The angles PHI0 and PHI1 are generally expressed as PHI0 = arcsin(U0.n R ) and PHI1=arcsin(U1.n R ) It is thus sufficient to define U1 as the vector associated with each sound reflection case. This way, the position of each of these obstacles can be determined for localization by acoustic detection, for applications in augmented reality or robotics.

[0093] Attitude of the reflecting surface n R If is initially unknown, then estimates of the wavefront parameters associated with at least two source positions where reflections from the same reflecting surface are detected, observed at different times, allow a complete estimation of its pose. Thus, a first set of parameters (U0, U1, TAU1) and at least a second set (U0', U1', TAU1') are obtained. If U0 and U1 define a plane perpendicular to the plane R, then their vector product defines the axis of this plane R, and the same applies to the vector product subtracted from U0' and U'1.

[0094] The product of these vectors (which are not collinear) defines the pose of the surface R.

[0095] However, one limitation of a model with only two interfering waves (direct and reflected) is that it can be difficult to distinguish between the various first reflections at the barrier. Furthermore, as additional reflections are introduced, the spatial-frequency behavior of the velocity vector quickly becomes more complex. Indeed, the paths of the real and imaginary parts may non-trivially vary along several axes. - In parallel planes for the direct wave and for two reflected sounds, - or, in general, throughout space, To be combined.

[0096] When several reflecting surfaces are considered, their complex spatial frequency distribution makes determining the parameters of the model very complicated.

[0097] One solution to this problem is to perform a time-frequency analysis with greater temporal resolution (i.e., a shorter time window) in order to have the opportunity to identify the appearance of a simpler acoustic mix at the onset of a certain amplitude (temporal, onset signal), i.e., to reduce the number of reflections interfering with the direct sound in the mix at that frame. However, in some situations, the delays associated with a series of reflections may be too close together to isolate the effect of the first reflection in interfering with the direct sound.

[0098] Therefore, the following process is proposed, which allows for easy separation and characterization of the effects of multiple interferences. The first step involves converting the velocity vector fingerprint into the time domain (i.e., "TDVV" for "Time-Domain Velocity Vector") using the inverse Fourier transform shown in Equation 31. This has the effect of reducing the effects of frequency cyclicity, which manifests as a complex itinerary of velocity vectors associated with an axis, into sparser, and therefore more easily analyzable, data. Indeed, such a transformation results in a series of peaks appearing at regular time intervals, making the most prominent peaks easier to detect and extract (see, for example, Figure 5B).

[0099] A notable property is that by construction (by inverse Fourier transform), the vector at t = 0 is equal in the frequency domain to the mean of the velocity vector (the mean of the real part, if only the positive frequency half-spectrum is considered). Such an observation is relevant for estimating the main DoA U0.

[0100] Starting from the frequency model of the velocity vectors for two interfering waves (direct sound and reflected sound), the denominator can be reformulated, as useful, by the Taylor expansion of Equation 32. Under the conditions for x and gamma1 given by Equation 32, the (conventional) velocity vector Equation 33 is reached. Under the condition that the amplitude of the reflected sound is smaller than that of the direct sound (g1 < g0 = 1, which is generally at the start of the sound), the inverse Fourier transform of this equation converges and is formulated as Equation 34, where the first peak is identified at t = 0, U0 (the direction of the direct sound) is obtained, and the characteristics of the peak series of the interference between the reflected sound and the direct sound are obtained.

[0101] These peaks are located at multiple times t = kTAU1 (non-zero integer k > 0) of the delay TAU1, and the amplitude in the norm decreases exponentially (following the gain g1). By using the conventional velocity vector, these are all associated with the direction that is on the same straight line as the difference U0 - U1. For this reason, they are orthogonal to the intermediate plane of these two vectors and alternate in direction (sign). The advantage of converting the velocity vector into the time domain is that the desired parameters are sparse and are expressed almost directly (Figure 2).

[0102] Thus, in addition to the main DoA U0, the following can be determined. - The delay TAU1, which may also be for some distinct walls, - Next, the vector that is on the same straight line as U0 - U1 and is normalized to the unit vector n, for example in Equation 41, and is available (given by the conventional velocity vector) for the following, - Assume U1 to be symmetric to U0 with respect to the intermediate plane, - Optionally, the attenuation parameter g1 (which can be modified by the time-frequency analysis parameters, particularly by the form of the analysis window and by its time arrangement based on the observed acoustic event. Therefore, the estimation of this parameter is less useful in the application scenarios considered here).

[0103] By observing multiple of the following time peaks, it is possible to determine whether they are substantially consistent as a series (multiple delays TAU1, multiple delays TAU2, etc.), thereby characterizing the same interference, otherwise it is necessary to identify, for example, that there are multiple reflections.

[0104] Below, we consider the "favorable condition" case. If the sum of gammas over N in Equation 35 is less than 1, and there are N reflections, we apply the Taylor expansion to obtain the conventional velocity vector according to Equation 35. The Taylor series transforms the denominator of the first equation, and can be rewritten using the polynomial law in Equation 36. The conventional velocity vector V model equation can be recognized as a "cross series" expressed as the SC term in Equation 37, as a sum of several series. For the general velocity vector V', we found a slightly different equation in Equation B2 at the end of the appendix. This equation is also expressed as Equation 35b because it corresponds to Equation 35 but is given for the general velocity vector. Furthermore, at the end of the appendix, an equation specific to the general velocity vector V' is given, and by adding a "b" after the equation number, it corresponds to the equation specifically for the conventional velocity vector (Equation xxb).

[0105] Under the condition of Equation 38 for the conventional velocity vector and for any frequency f (Equation B3 = 38b for the general velocity vector), the following time series Equation 39 (Equation B4 = 39b for the general velocity vector) combined with delay SARC is deduced by inverse Fourier transform. The first peak at t = 0, which gives U0 (direction of the direct sound), is identified, and for each reflection, a series of peaks characteristic of the interference of the reflection with the direct sound is identified. For example, in Figure 2, these peaks are located at successive positive time coordinates TAU, 2TAU, 3TAU, etc., which are multiples of the delay TAU between the reflection at the wall and the straight path.

[0106] Then a series of features emerges (at larger time coordinates) of interference between some reflected sound from some walls and the direct sound, where the delay is another combination (in positive multiples) of various delays.

[0107] Indeed, Figure 2 shows such a sequence for the simple case of two reflections interfering with a direct sound. Each marker (circle, cross, diamond) indicates on the ordinate the contribution of a vector U0, U1, U2 (characteristic of the direct sound, the first reflection, and the second reflection, respectively) to the time fingerprint TDVV as a function of the time coordinate. It can thus be seen that the reception of the direct sound is characterized by a first peak with time zero and amplitude 1, indicated by a circle. The interference of the first reflection (with delay TAU1) with the straight path gives rise to a sequence of first peaks at TAU1, 2xTAU1, 3xTAU1, etc., marked with a cross at one end and a circle at the other (top and bottom). The interference of the second reflection (with delay TAU2) with the straight path results in a second series of peaks at TAU2, 2xTAU2, 3xTAU2, etc., marked with a diamond at one end and a circle at the other. This then leads to the "crossover series", i.e., the interference of reflections with reflections (first delay: TAU1+TAU2, then 2TAU1+TAU2, then TAU1+2TAU2, etc.). The crossover series, while achievable in the general case, is lengthy to write out and, for the sake of brevity, is not written explicitly here, especially since it is not necessary for its use in estimating the relevant parameters in the process presented here.

[0108] In the following, we describe the analysis of the temporal fingerprint by successively estimating the parameters.

[0109] The estimation of the model parameters based on the calculated time series is performed in the same way as in the single reflection case described above. First, let us consider the most general situation (excluding the special cases we will discuss later). This corresponds to the preferred case where the delays are not "overlapping" and the above-mentioned sequences do not coincide in time. That is, each identifiable peak belongs to only one of them. Therefore, by designating the time peaks as increasing delays from t=0, a new peak detected with a delay TAUnew either belongs to an already identified sequence or defines the beginning of a new sequence. Indeed, considering a set of delays characterized by already identified reflections, the first case is detected if a positive or partially zero integer k results in TAUnew according to Equation 40. Otherwise, the second case occurs, and the identified set of reflections is augmented by introducing a new delay TAUN+1, which is associated with an estimable direction in the same way as described for the single reflection case.

[0110] In practice, it will not be necessary to attempt to explain many of the time peaks. We will limit ourselves to the first peak observed, especially since it is the most easily detectable, since its amplitude (i.e., absolute magnitude) is greater than: Thus, situations in which the delays are common multiples, but of higher (or lower) rank Ki;Kj, as a function of amplitude, can be analyzed by the above process.

[0111] As long as the sum of the absolute values ​​of the implicit gains gn(n>0) is less than 1 (Equation 38 for conventional velocity vectors), the inverse Fourier transform (Equation 31) yields a unidirectional time fingerprint that evolves in positive time.

[0112] On the other hand, if the sum of the absolute values ​​of the implicit gains gn (n>0) is greater than 1, the inverse Fourier transform results in a "bidirectional" time fingerprint TDVV, with the sequence generally evolving toward both positive and negative time (the upper curve in FIG. 5B for illustration). This situation, where one or more reflection gains are greater than 1, can occur, for example, when the direct wave has an amplitude less than the sum of the amplitudes of the waves resulting from the reflections at one or more barriers. In this "unfavorable case," the main peak at time zero no longer corresponds strictly to the vector u0, but to a combination of the latter and a somewhat large proportion of the vector indicating the direction of the reflections. This leads to positioning bias (in the "estimated DoA"). Another symptom is that the main peak generally has a norm different from 1 and, more often, is less than 1.

[0113] The present invention proposes a method that is particularly stable in this situation: it is proposed to adjust the velocity vector formula by providing a spatial separation towards DoA with the D component appearing in the denominator instead of the usual omnidirectional component W.

[0114] By providing a direct approach to the reference component D, the relative attenuation associated with each reflection of index n is assigned to the denominator by the coefficient BETAn, while the overall coefficient NU0 is calculated (Eq. B1 in the Appendix). This leads to the formula for the general velocity vector given by Eq. B2 = 35b for the model of N reflections, which was already introduced for the general velocity vector case. Note that the condition for the Taylor series expansion, as presented on the right-hand side of the equation, is given by Eq. B3 = 38b. With the additional attenuation coefficient BETAn, this condition proves to be more easily satisfied in many situations. This condition, and the resulting model-based index, confirms the causal nature of the entire time series. Under this condition, which is satisfied for all frequencies, the general velocity vector model in the time domain is specified in the form of Eq. B4 = 39b, which (as in the case of the conventional velocity vector in Eq. 39) - the first peak at t=0, whose direction gives the DoA, and the vector U0 is obtained by normalizing equation B5, - the same number of columns as the reflected sounds, each associated with the interference of the reflected sound with the direct sound, for which the values ​​of the vectors observed at regular intervals TAUn appear in equation B6, - indicates a sequence associated with a delay, denoted as SARC, that is not used in the subsequent estimation procedure.

[0115] Starting with equation B6 at the end of the appendix, we obtain specific relationships between successive vector columns, most notably the first two vectors V'(TAUn) and V'(2.TAUn). This gives us equation B7, which shows a coefficient (-Gn / BETAn), denoted "RHO," and equation B8, which provides an estimate as the scalar product of the first two vectors V'(TAUn) and V'(2.TAUn) in the same column, divided by the square of the norm of the first one. By inserting the RHO coefficient back into equation B6, we can rearrange the resulting equation B9 to obtain equation B10. The right-hand side of this equation denotes the vector Un (specifically, the vector U1, if we consider the first reflection and its associated column) and is affected by the coefficient NU0 / Gn, which is positive (except in perhaps rare situations such as reflections with phase inversion) and is therefore obtained by normalizing the left-hand side V'(TAUn)-RHO.V'(0).

[0116] Furthermore, the overall coefficient NU0 is more likely to integrate other influencing coefficients than the reference directivity, resulting in, for example, an overall amplitude reduction, which can occur by limiting the frequency bandwidth of the signal source and / or by partially masking noise (although the latter effect is generally more complex to model). Interestingly, ultimately, the direction of the vector U1 (or more generally Un) can be estimated as well, and is the reason for this overall amplitude reduction NU0.

[0117] This estimation aspect also applies to conventional velocity vectors (in which case it is simply necessary to consider BETAn=1).

[0118] In particular, practical examples of embodiments in which general velocity vectors are used to determine parameters such as DoA are described below.

[0119] In the present embodiment, which will be described with illustrative reference to Figures 6A to 6D and Figure 7, a first estimate of the delay (step S71 in Figure 7) is obtained, and a conventional velocity vector V(f)=1 / W(f)[X(f),Y(f),Z(f)] T is now calculated "normally", for example, for the first order Ambisonic component.

[0120] In step S721, the above calculations are performed based on the frequency expression of the conventional velocity vector V(f) up to the estimation of the time expression of the conventional velocity vector V(t).

[0121] In step S731, a conventional analysis of the time expression of the velocity vector V(t) is performed as a peak time series. This analysis, in particular, determines whether an efficiently (low interference) identifiable, unidirectional time series evolves only in positive time (as a truly "causal response"), as described in Appendix Equations 39 and 40.

[0122] In a time-domain analysis of the conventional velocity vector V(t) equation, negative time peak structures (e.g., energy or amplitude higher than a selected threshold THR) are typically found, such as the peak on the negative abscissa in FIG. 6A. This means that they cannot be properly identified in Equation 39, and therefore the DoA estimate given by the peak V(t=0) is biased. This is also an indication that the convergence condition of the Taylor series leading to Equation 39 is not satisfied, and therefore the amount of indirect waves mixed with the direct sound in the calculation of the denominator of V(f) is too large relative to the direct sound. In the proposed improvement, the proportion of these indirect waves is reduced by a spatial filter. This means that the directivity (acting in the denominator) in the estimation of the velocity vector V(f) needs to be improved.

[0123] A spatial filter is then applied to the acquired Ambisonic data (step S71) to beamform in the direction (DoA) estimated in step S751 from the previously acquired velocity vector V(t==0). This first estimate, while possibly in error, provides an approximation. While admittedly coarse, this approximation is sufficient to point the beam toward the direct sound source and reduce reflected sounds coming from more distant angles.

[0124] A modified velocity vector V'(f), and in the time domain V'(t) (step S781), is calculated based on these filtered Ambisonic data.

[0125] Test S732 determines whether the time expression of the corrected velocity vector V'(t) has any remaining peaks with time coordinates less than 0. Even if an improvement over the preceding Fig. 6A is observable, it can be determined whether signal structures at negative time coordinates (e.g., estimated as energy (shown as "NRJ" in Fig. 7), e.g., given by the integration of the signal at negative time) remain greater than a threshold THR, as shown by way of example in Fig. 6B.

[0126] In this case, the method repeats (S752) by obtaining a coarse estimate of the previously obtained DoA(n), determining a reference component D(f) (denoted as D(n) for the nth iteration of the method in step S762) whose directivity indicates the direction of the direct sound, with higher resolution than the estimate D(n-1) from the previous iteration, and replacing D(n-1) in the estimation of the generalized velocity vector V'(f) (S772), obtaining V'(t) in step S782. In this way, the directivity of the reference component "captures" the estimated direction of the direct sound with higher resolution than the reference component from the previous iteration. In this embodiment, the Ambisonic order does not need to exceed 1, and both the pose and shape of the directivity can be adjusted to better capture the direct sound, for example, to further eliminate certain reflected sounds.

[0127] In this manner, the method may be repeated until the amplitude or energy of the peak at negative time is less than a selected threshold THR, as shown in FIG. 6C.

[0128] In this way, increasing separation toward the direct sound is continually propagated to the component in the denominator of the velocity vector (traditional and common) in a linear equation as the iterations proceed. Figure 6D thus transitions from an omnidirectional W component in a spherical (light gray) shape to a more highly separated component D(1), in this example a supercardioid shape in dark gray, and then to a narrower supercardioid shape D(2), in very dark gray.

[0129] Returning to the method shown in FIG. 7 and explaining in more detail, the first step S71 begins with constructing a conventional velocity vector V(f) in the frequency domain with the omnidirectional component W(f) in the denominator. In step S721, the time-domain expression V(t) is estimated. Then, in test S731, a signal structure is identified, showing the time expression of the conventional velocity vector V(t) with a peak, and if the energy NRJ of this signal structure at negative time coordinates (t<0) remains below a fixed threshold THR (arrow NOK), and if the current acoustic situation allows for an unbiased DoA to be extracted directly from the conventional velocity vector, then the parameters DoA, U0, U1, etc. can be determined in step S741 as described above. Otherwise (arrow OK leaving test S731), the direct estimation of DoA from the conventional velocity vector is biased, and at least a first iteration (n=1) must be performed to adjust the velocity vector to determine a general velocity vector, as described below.

[0130] From this DoA measurement, even if biased (obtained in step S751), a reference component D(1) is estimated in the frequency domain in step S761, which in turn replaces the omnidirectional component W(f) in the expression of the "general" velocity vector V'(f) in step S771. In step S781, the time expression of the general velocity V'(t) is estimated in test S732 to determine whether significant energy (above the threshold THR) remains in the signal structure of this expression V'(t) in the negative time coordinate. If not (arrow NOK leaving test S732), the process stops in this first iteration by providing the parameter DoA etc. and proceeds to step S742. Otherwise, the process is repeated by updating the iteration index n in step S791 (where the step denoted S79x determines whether the iteration is based on the index n (step S793) or ends steps S792-S794).

[0131] Based on the coarse DoA estimated in the previous iteration (step S752), as described above, a new reference component D(n) is estimated in step S762, replacing the previous reference component D(n-1) in the denominator of the generalized velocity vector V'(f) in step S772. From this new expression of the generalized velocity vector V'(f) in the frequency domain, its expression V'(t) in the time domain is determined in step S782. A comparison of the signal structure (energy relative to a threshold THR) is repeated in test S733 to determine whether the new DoA that can be estimated from this is biased. If not (arrow NOK leaving test S733), the parameters, especially the DoA, can be obtained in step S743 after three iterations, as shown in the example of Figures 6A-6C. Otherwise (arrow OK leaving test S733), the process has to be repeated again from step S752 using the last estimated DoA, albeit coarse and possibly biased.

[0132] A possible example of calculating the reference component D(f) from a previously estimated DoA is described below. In the form found in the equations in the appendix, the component D Θ Since (f) takes the following matrix of Ambisonic components (or vector D Θ (from a sum weighted by

[0133] D Θ (f)=D Θ .B(f), where: - B(f) is a vector of signals describing the Ambisonic field in the frequency domain, e.g.

number

number

number

number

number

number

number

number

number

[0134] Depending on the first order, spherical functions

number

[0135] In the specific case where the directivity is formed axially symmetrically, this is D Θ =Y(Θ).Diag(g beamshape ) The format is:

[0136] And the gain

number

number

number

[0137] Diagonal coefficient g beamshape takes into account the following: - On the one hand, the choice of Ambisonic coding and therefore spherical harmonics

number

[0138] For these aspects, it is useful to refer to the paper by Jerome Daniel (2000), in particular pages 182-186 and Figure 3.14, where the tools proposed for spatial decoding are directly applicable to directional construction from Ambisonic signals, here shown as D-reference components.

[0139] In addition, following this format, gbeamshape A gain factor can be defined that provides separability for the above component D(f) as a function of

[0140] 7, acoustic conditions may prevent a good DoA estimate from being obtained because no form for the generalized velocity vector V'(t) can be found in the time domain from a good "causal" perspective. Referring again to FIG. 7, a termination criterion may be added to exit the iteration of processing if, for example, the shape of the signal representing the generalized velocity vector V'(t) in the time domain cannot be further improved. Thus, at the end of the previous iteration, test S733 determines that the energy of this signal is greater than a threshold THR at negative time coordinates (t<0) and the iteration of processing does not improve (i.e., does not further improve) the estimate of the generalized velocity vector V'(t), i.e., if both of the following can be indicated: - signal energy exceeding the threshold THR at negative times, and - the signal energy that does not decrease when going from one iteration n-1 to the next iteration n (arrow NOK leaving test S792), The process iteration may stop at step S794.

[0141] Otherwise (arrow OK leaving test S792), the process is ready for the next iteration and begins in step S793 by incrementing the iteration counter n.

[0142] In order to minimize the occurrence in such "worst case" acoustic situations where convergence on an unbiased DoA is not reached and processing is forced to stop at step S794, the teachings of the above-mentioned document FR1911723 can be adopted, for example, for a solution that makes it possible to pick the best frame to increase the chances of an unbiased decision (e.g., the start of a signal frame).

[0143] Indeed, as described in document FR1911723, the relative importance of this problem makes it possible to assess to what extent the vector U0 is likely to provide a good estimate (with low bias) of the DoA, thus providing a confidence factor for the estimation and allowing the estimate obtained for a certain frame to be retained preferentially. If the risk of estimation bias proves to be excessive, a frame that faces this problem the least can be selected, as will be explained below with reference to Figures 3A and 3B.

[0144] The embodiment described below can then be applied to the estimation of conventional velocity vectors, in particular, for example during the first iteration of the process described above with reference to Figure 7. It will be assumed below that such a process is applied to conventional velocity vectors, as already explained in document FR1911723.

[0145] Thus, a frequency analysis of a time subframe can proceed to observe the first peak for a given space. The frame with the onset of the signal (rising energy, transient, etc.) allows observing a sound mix containing only the fastest wavefronts: the direct sound and one or more reflected sounds (so that the "gamma sum" mentioned above remains less than 1, according to Equation 38).

[0146] For the frame containing the onset of the signal, the time window can be adjusted (possibly dynamically) for frequency analysis, e.g., by making it asymmetric and generally decreasing in shape, so that the "roughness" of the window gives more weight to the rising edge (onset, transient) of the signal, and thus gives progressively less weight to the direct sound (e.g., approximately exponentially, but this is not required). In this way, the amplitude of the later wavefronts is artificially reduced compared to the earlier wavefronts, reaching the convergence condition of the Taylor series and ensuring a unidirectional time evolution.

[0147] An example of the exponential decay type of time windowing is shown below. This example is applied to the analyzed signal in order to derive the resulting time fingerprint in a favorable case without substantial bias in estimating the wave's direction of arrival. For convenience, this operation is performed effective from time t0, designated as the zero time point. This time preferably corresponds to the beginning of the signal, which begins with silence. As in Equation 42, ALPHA>0, and by again integrating the convolution form involving s(t) and h(t), the form of Equation 43 is obtained.

[0148] Then, equation 44 properly contains an exponential function to justify this choice, and we obtain the form given in equation 45, which means establishing equation 46.

[0149] Therefore, modeling the impulse response due to a set of reflected sounds added to a direct sound gives Equation 47:

[0150] In this way, if the sum of the gammas is greater than or equal to 1 (sometimes referred to as a "bidirectional series"), it is always possible to determine the attenuation coefficients ALPHA such that the sum of the "adapted" gains (Equation 48) is less than 1.

[0151] It is then observed that the temporal fingerprint is indeed unidirectional, with peaks evident only at positive times after application of the decreasing exponential window (bottom of Figure 5B). In fact, it is further observed that the energy of the observed signal decays exponentially very quickly, and the numerical impact of signal truncation on the estimate becomes completely negligible beyond relatively short truncation times. In other words, the benefits of long-term analysis of both the entire excitation signal and its reverberations are realized at shorter times. Indeed, the observed "TDVV" conforms to the interference model without errors due to signal dynamics. Such window weighting therefore has the dual property of ideally yielding a usable temporal fingerprint.

[0152] In practice, it is necessary to determine the attenuation at ALPHA without knowing the amplitude of the reflected sound in advance, preferably a compromise is sought between a value low enough to ensure unidirectionality of the time fingerprint, but not too low to avoid reducing the chance of detecting and estimating the indirect wave. For example, this value can be chosen to be a time t that is physically representative of the measured phenomenon. EXP (usually 5ms) attenuation coefficient a EXP For ALPHA=-(log a EXP ) / t EXP can be determined as follows:

[0153] An iterative process (e.g., by bisection) can be implemented to adjust the attenuation values: if a threshold attenuation value is exceeded and the acquired temporal fingerprint is detected in both directions, then the analysis is repeated with stronger attenuation, with a vector U0 that is therefore possibly biased; otherwise, at least the estimate U0 is applied; if the following peaks are not well distinguishable (due to attenuation reduction), the analysis is repeated with an attenuation intermediate between the two preceding ones, and so on, if necessary, until the vector U1 can be estimated.

[0154] Nevertheless, the exponentially decreasing window approach can be sensitive to interference, especially at the beginning of the windowing process, where interference is significantly amplified. At the beginning of the windowing process, if the start is very recent, interference other than noise may simply be the echo of the source itself. To reduce such interference, a noise reduction process can be implemented.

[0155] In general, time windows of various shapes and / or sizes may be provided, and there may be overlap between the windows to maximize the chances of obtaining a "good fingerprint."

[0156] The initial DFT size is chosen to be generally larger than this analysis window.

[0157] This is of course the case in the context of digital audio signal processing, where signals are sampled at a given sampling frequency in the form of successive blocks of samples (or "frames").

[0158] Pre-processing and time-frequency denoising can also optionally be provided for onset, transient etc. detection, for example by defining a mask (a time-frequency filter, which may be binary) to avoid introducing other environmental elements and / or to diffuse the field source into the interference fingerprint. The impulse response of the mask (result of the inverse transform) must be calculated to verify its effect on the analysis of the peaks. Alternatively, a weighted average of frequency fingerprints, possibly corresponding to similar interfering mixtures, can be integrated into the frequency weighting of the fingerprint of the stored frame, to be subsequently calculated (typically by making sure that the source in question has not moved at the time of signal onset, which can be inferred from a delay estimate).

[0159] And thus proceed to extract and observe the peaks, for example by giving the norm |V(t)|: maximum peak, and then TAU1 (in general), etc.

[0160] Then we proceed to analyze the temporal fingerprint, which (according to {tau_n} and V(sum(k_n.tau_n)))) is: - whether there is a time loop (a kind of circular "aliasing") due to the choice of FFT on too-short temporal support, - Is there a progressive unidirectional series, or conversely, a bidirectional series? Or, for certain series without noticeable decay (where the sum of the gains (gn) remains close to 1), or for retrograde series (at least for implicit gains g_n>1), This is done by detecting

[0161] and, - assign and store a "good frame" or "good fingerprint" score (allowing for a reliable estimate, unidirectional and therefore possibly free of DoA bias); - Obtain an estimate (Un) - If necessary, adjust the upstream analysis by selecting an appropriate time window.

[0162] The analysis of the temporal fingerprint has been described above, but the frequency analysis can be performed more simply as follows.

[0163] It can be easily seen mathematically that the peak at time zero is, by construction, equal to the average of the velocity vector over the complete spectrum (which has no real part due to the symmetry of the Hamiltonian), or even its real part if only positive frequencies are considered. If one is only interested in the direct sound, it can be deduced that there is no point in computing the inverse transform of the FDVV to obtain an estimate of the DoA. However, a time-based estimate of the TDVV makes it possible to detect whether this DoA is reliable (a criterion that evolves with positive and progressing time).

[0164] This preferred case is more likely to be observed at the onset of the original signal, unless the mixture is still very complex, which is generally sufficient to obtain measurements at these times.

[0165] Furthermore, in practice, the frequency and time fingerprints of VVs are not always distinguishable with ideal models of interfering wave mixtures. The original signal may not fully or always excite a significant range of frequencies at each key point in time due to a lack of transmit power, or may take into account the concurrency of other components of the captured sound field (insufficient SNR or SIR). This can lead to a somewhat diffuse background sound (other sound sources), microphone noise.

[0166] Processing can then be applied according to at least one of the following, or a combination of several: - selection of time-frequency samples with signal onset detection according to an improved algorithm; - over several frames (e.g., possibly dynamically, |W(f)| 2 and the mean of V(f) weighted by a forgetting factor) velocity vector smoothing, - For the selection of the signal start frame, |W(f)| 2 Calculate the average of V(f) weighted by (for the same extracted delay) to supplement the frequency fingerprint and strengthen the time fingerprint. For efficiency in calculation, it may also be advisable to perform the calculation of TDVV, or upstream FDVV, only for frames that are detected as being more consistent in information, e.g., signal start frames, in situations where this information can be detected by simple processing, in which case it is still advantageous to position the analysis window at the rising edge of the signal.

[0167] For good estimation of non-integer delays (fractional delays and multiples thereof in the time series), peak estimation by sample-to-sample interpolation and / or local frequency analysis (by isolating peaks in time-constrained neighborhoods) can be considered, and the delay is fine-tuned based on the phase response.

[0168] The pre-selection of the time peak can be done based on the current estimate of the characteristic delay of the series.

[0169] Thus, the steps performed are summarized in a possible exemplary embodiment as shown in Figures 3A and 3B. In step S1, the Fourier transform (from time to frequency) of the Ambisonic signal is calculated, which is done in the form of a series of "frames" (series of blocks of samples). For each transformed frame k (step S2), a dynamic mask can be applied to some frequency bands whose signal-to-noise ratio is below a threshold (some frequency bands may actually be noisy, e.g., noise inherent in the microphone or otherwise impairs the usefulness of the signal acquired in this frequency band). In particular, a search for noise by frequency band is performed preferentially for the "omnidirectional" component W in step S3, and frequency bands modified by noise (e.g., exceeding a threshold such as SNR<0 dB) are masked (i.e., set to zero) in step S4.

[0170] Then, in step S5, the velocity vector V(f) is calculated in the frequency domain, for example, according to Equation 6 (or in the form of Equation 11, Equation 18 or Equation 20).

[0171] Here, in one exemplary embodiment, weights q(f), calculated as described below, are applied to assign some importance to frequency band f. Such an implementation allows the velocity vector V(f) to be represented in frequency bands where evolution is significant. To do this, optimal weights are iteratively calculated as a function of U and V(f). Thus, returning to the algorithmic process of FIG. 3A, in step S6, various weights q(f) are set to 1. In step S7, the weights q(f) applied to V(f) for each band are applied such that V = q(f)V(f). In step S8, U is determined for each frame k, and U0(k)=E(Re(Vbar(f))), where E(x) is, for example, the expectation of x and is therefore equivalent to the average over all frequencies of the real part of the estimated velocity vector Vbar(f).

[0172] This first estimate U(k) is naturally coarse. It is iteratively adjusted using Equation 49 based on the imaginary part of vector V(f) by calculating weights for U(k) in the preceding discrimination, where vector m is a unit vector perpendicular to the plane defined by vector U and normal to the wall (e.g., direction z in FIG. 1). Vector m is also iteratively estimated as a function of U in step S9, and weights are calculated using Equation 49 in step S10. The resulting weights are applied in step S7, and the estimate of U is adjusted until convergence is reached at the output of test S11. At this stage, U(k) has been estimated for various frames.

[0173] From this, U1 can be deduced by a relationship such as Equation 41 above. In the variant described here, U1 is determined by Equations 50 to 52. These equations are derived by previously applying an inverse IDFT (frequency to time) transform to the vector Vbar(f) obtained in step S7 in step S12 to obtain a time-based representation V(t) of the velocity vector. With this embodiment, various delays TAU1, TAU2, etc. can be determined for each different reflecting surface, as can be seen with reference to FIG. 2. The first delay TAU1 is determined because it is the first peak V(t) at a time following the moment of reception of the straight path. Thus, in Equation 51, tmax(k) is the instant that maximizes the absolute value of V(t) calculated for frame k.

[0174] Test S13 verifies for each frame that the absolute value of V(t=0) is indeed greater than V(t) for t>0. Frames that do not satisfy this condition are discarded in step S14. The various delays TAU1 and TAU2 are then determined in step S15 (by removing the one corresponding to the delay TAU1 from the absolute value of V(t)k compared in equation 51), etc. The delay TAUm is given by dividing the component tmax obtained at each iteration m by the sampling frequency fs according to equation 52, taking into account that the times t and tmax(k) are first expressed in terms of sample index (time zero is used as the reference for index zero). The vectors U1, U2, etc. can then be calculated according to equation 50.

[0175] Other parameters can also be determined, in particular d0 given by equation 28, in step S16 (by then verifying the consistency of conventional indoor data, such as d0min=0 and d0max=5m, in test S17; if this is not the case, the frame contains an error and can be discarded in step S14).

[0176] Step S18 consists of selecting a "good" frame that represents the onset of the sound at the first reflection. The criterion D(k) for selecting such a frame can be described, for example, by Equation 53: where C(f)i(k) specifies the magnitude (absolute value of amplitude) detected in Ambisonic channel i at time-frequency sample (t,f) resulting from the first transformation (time to frequency) of frame k. Epsilon specifies a non-zero positive value to avoid zero denominator due to no signal. F specifies the total number of frequency subbands used.

[0177] In this way, among the criteria of all frames D(k), only frames whose criteria D(k) calculated from Equation 53 is not smaller than 90% of the maximum value Dmax obtained in step S21 can be retained in step S22.

[0178] Thus, in step S18, the value of D(k) is calculated for every frame, and in step S19, the process provides U0(k), d0(k), and D(k) for the various frames. In step S20, the values ​​of D(k) are collected to identify the highest value in step S21 and discard frames with D(k) values ​​less than 0.9Dmax in step S22.

[0179] Finally, in step S23, the vector U0 retained here is preferably the median (rather than the average) of the vectors U0 of the various frames retained, and the distance d0 retained is the median of the distances d0 of the various frames retained.

[0180] The invention is of course not limited to the embodiments described above by way of example, and this also applies to other variants.

[0181] An example application to the processing of first order Ambisonic (FOA) signals has been described above. The order may be increased to enrich the spatial resolution.

[0182] While a first-order Ambisonic representation has been described, higher orders are also possible. In such cases, the computational complexity of the velocity vector increases with the ratio of the higher-order directional components to the component W(f), and the vector Un increases implicitly with increasing order. The increased dimensionality (beyond three), and the resulting spatial resolution, allows for better discrimination between the vectors U0, U1, ..., Un, and facilitates the detection of the peak V(k*TAUn) proportional to (U0-Un) for the temporal fingerprint, even when the vectors U0 and Un are angularly close, as occurs with grazing reflections (e.g., when the source is far and / or close to the ground). This therefore allows for more accurate estimation of the desired parameters U0, U1, d0, etc. It is also noted that only the three first-order components (X, Y, Z) are retained in the numerator here, so that the higher-order components are available to construct the reference component in the denominator, making it independent. In all cases (regardless of the denominator), it may be considered to improve the above process (and the process shown in FR1911723) by adding higher order components to the numerator, thus increasing the dimension of the velocity vector and making the peaks more distinguishable, especially in the time domain.

[0183] More generally, velocity vectors can be replaced by ratios between components of a spatial acoustic representation that are "simultaneous" in the frequency domain, and can function in a coordinate system characteristic of the spatial representation described above.

[0184] To mitigate the case of multiple source instances, more generally, artificial intelligence methods, including neural networks, can be used to calculate TDVV. With certain possible learning strategies (e.g., on fingerprints from models or windowed SRIR, and not necessarily from the original signal), the network can learn to use a sequence of frames to detect and estimate a given indoor situation.

[0185] Furthermore, the possibility of first estimating a conventional velocity vector V(f) to determine a first rough estimate of the DoA, and then adjusting the estimate of the general velocity vector V'(f) based on this rough estimate to obtain a more accurate DoA, was described above with reference to FIGS. 6A-6D and 7. It should be understood that this is merely one possible implementation example. For example, in one variant, the space can be directly divided into several parts, and thus the components in the denominator D(f) in each of these parts are given a directionality. The DoA calculation attempts to converge by an iterative method (of the type shown in FIG. 7 starting from step S761, where D(f) can be simply calculated as a function of the angular part), and these iterative methods are in particular performed in parallel for each part. This is one possible exemplary implementation example. Alternatively, the "best" (or several best) angular parts can be selected according to the above-mentioned criterion of causal model effectiveness, and the estimate optimized for that direction or directions is selected and included in the denominator as a variant of the form. Indeed, more generally, one can further foresee evaluating, in a first step and / or during subsequent steps, multiple general velocity vectors each associated with a differently shaped directivity (in the denominator).

[0186] More generally, the search for a beamforming solution that provides reliable direction of arrival (DoA) of sound can be explored as a general optimization problem, within which various strategies can be required. - a minimization criterion (function to be minimized) that expresses / predicts the effectiveness of the causal model. In a somewhat simple and therefore improvable way, in the above description, this focuses on searching for the relative energy of the signal in the negative time coordinate. The parameters to be optimized are the beamforming parameters, for example the coefficients of the matrix D contained in equations A1 and A4, or the beamshape parameters g in equation EQ.A5, formulated as theta (Θ) direction and if axial symmetry is chosen, i.e., restricted to first order. beamshape It is formulated as follows. - The first set of parameters (or sets of parameters) is usually "omnidirectional" (g beamshape = [1 0...]), or a preferred directionality stored in previous use, or other multiple directionality referring to a set of directions representing space, and also representing various shapes of directionality. - The principle of adjusting (in an iterative process) the tested parameters, since usually repositioning the acoustic beam at the latest estimated DoA is not always a sufficiently stable choice. Rather than stopping the algorithm due to lack of improvement, it is necessary to start again from one of the stored situations (perhaps the best from the point of view of the minimization criterion) and adjust the parameters along other axes (for example, in the form of directivity parameters) or other combinations of axes.

[0187] More generally, conventional techniques including, for example, stochastic gradient or batch optimization are considered, but the large number of iterative processes can be quite costly.

[0188] Note that, unlike common optimization tasks, the final objective parameters (typically an Un vector) are not directly optimized but result from it. Note that there are potentially many different sets of beamforming parameters that are all "optimal" in the sense that they result in adherence to a causal model, allowing the same set of parameters Un to be inferred, possibly accurately.

[0189] Addendum

[0190]

number

[0191]

number

[0192]

number

[0193]

number

[0194]

number

[0195]

number

[0196]

number

[0197]

number

[0198]

number

[0199]

number

[0200]

number

[0201]

number

[0202]

number

[0203] Equation 13 g0=1;τ0=0

[0204]

number

[0205] Equation 15 γ0=1

[0206]

number

[0207]

number

[0208]

number

[0209]

number

[0210]

number

[0211]

number

[0212]

number

[0213]

number

[0214]

number

[0215]

number

[0216] Equation 27 d1-d0=τ1c

[0217]

number

[0218]

number

[0219]

number

[0220]

number

[0221]

number

[0222]

number

[0223]

number

[0224]

number

[0225]

number

[0226]

number

[0227]

number

[0228]

number

[0229]

number

[0230]

number

[0231]

number

[0232] formula 44 e -αt =e -α(t-τ) .e -ατ

[0233]

number

[0234]

number

[0235]

number

[0236]

number

[0237]

number

[0238]

number

[0239]

number

[0240]

number

[0241]

number

[0242] Formula A1 D Θ (f)=D Θ .B(f)

[0243]

number

[0244]

number

[0245]

number

[0246] Formula A5 D Θ =Y(Θ).Diag(g beamshape )

[0247]

number

[0248]

number

[0249]

number

[0250]

number

[0251]

number

[0252]

number

[0253]

number

[0254]

number

[0255]

number

[0256]

number

[0257]

number

[0258] MEM working memory PROC Processor OUT Output interface IN Input interface

Claims

1. 1. A method for processing an acoustic signal acquired by at least one microphone, comprising: To locate at least one sound source in a space having at least one wall, - a time-frequency transformation is applied to the acquired signal, - based on the acquired signals, a generalized velocity vector V'(f) is estimated from an equation for the velocity vector V(f) expressed in the frequency domain and in which a reference component D(f) other than the omnidirectional component W(f) appears in the denominator, the equation being a complex equation with real and imaginary parts, said generalized velocity vector V'(f) being: a first acoustic path, represented as a first vector U0, which is a straight line path between the sound source and the microphone; characterizing a combination of at least a second sound path caused by reflections at the wall and represented as a second vector U1; the second path has a first delay TAU1 at the microphone relative to the linear path; as a function of the delay TAU1, the first vector U0, and the second vector U1, * The direction of the straight line path (DoA); *The distance d0 from the sound source to the microphone; and a distance z0 from the sound source to the wall,

2. 2. The method of claim 1, comprising multiple iterations in which the generalized velocity vector V'(f) includes, at least in part, a reference component D(f) in its denominator determined based on an approximation of the direction of approach (DoA) obtained in a previous iteration.

3. a first iteration in which the velocity vector V(f) is used instead of the general velocity vector V'(f), and at the end of the first iteration, the velocity vector V(f) is expressed in the frequency domain and has the omnidirectional component W(f) in its denominator to determine at least a first approximation of the direction of approach (DoA); 3. The method of claim 2, wherein, for at least a second iteration following the first iteration, the generalized velocity vector V'(f) is used and estimated from an equation for the velocity vector V(f) in which the omnidirectional component W(f) is replaced by the reference component D(f) in the denominator, the reference component D(f) being more spatially resolved than the omnidirectional component W(f).

4. 4. The method of claim 3, wherein the reference component D(f) has a higher separation than the omnidirectional component W(f) in a direction corresponding to the first approximation of the direction of the straight path (DoA).

5. The method according to claim 2 , wherein the iterative process is repeated according to a causality criterion until convergence.

6. At each iteration, - an inverse frequency-to-time transform is also applied to the expression for the generalized velocity vector V'(f) in order to obtain, in the time domain, a series of peaks each associated with a reflection of sound from at least one wall in addition to the peaks associated with the arrival of sound emitted by the sound source along the linear path; - a new iteration is performed if a signal appears in the series of peaks whose time coordinate is less than that of the linear path and whose amplitude exceeds a selected threshold, The method of claim 5 , wherein the causality criterion is met if the amplitude of the signal is less than the threshold.

7. The iterative process - in a first case, the amplitude of the signal is less than a selected threshold; in a second case, where the repetition of the iterative process does not result in a significant decrease in the amplitude of the signal; The method of claim 5 or 6, wherein the method is terminated.

8. Following the second case, the following steps are carried out, the acquired signal being obtained in the form of successive frames of samples, each step comprising: For each frame, a score is estimated for the presence of a sound onset (equation 53) in that frame (S18); The method according to claim 7, further comprising a step (S22) in which frames with a score higher than the threshold are selected for processing the acoustic signal acquired within said frames.

9. The acquired signal is received by an Ambisonic microphone, and the velocity vector V(f) is represented in the frequency domain by first order Ambisonic components of the following kind of form: V(f)=1 / W(f)[X(f),Y(f),Z(f)] T , W(f) is the omnidirectional component, The generalized velocity vector V'(f) is represented in the frequency domain by first order Ambisonic components of the following kind of form: V(f)=1 / D(f)[X(f),Y(f),Z(f)] T , The method of claim 1 , wherein D(f) is the reference component other than the omnidirectional component.

10. 10. The method according to claim 1, wherein an estimate of the direction of the straight path equivalent to the first vector U0 is determined from the average value over a set of frequencies of the real part of the generalized velocity vector V'(f) expressed in the frequency domain (Equation 24).

11. - an inverse frequency-to-time transform is applied to said generalized velocity vector to represent it in the time domain V'(t), - after a time of said linear path, at least one maximum value is searched for in the expression of said general velocity vector V'(t)max as a function of time; 11. The method according to claim 1, wherein a first delay TAU1 is deduced therefrom, which corresponds to the time of the maximum value V'(t)max.

12. said second vector U1 is estimated as a function of the values ​​of the normalized velocity vector V' recorded at time indices t=0, TAU1 and 2×TAU1, so as to define a vector V1 as follows: V1=V'(TAU1)-((V'(TAU1).V'(2.TAU1)) / ||V'(TAU1)|| 2 )V'(0), The vector U1 is U1=V1 / ||V1|| The method of claim 11 , wherein the

13. the angles PHI0 and PHI1 of the first vector U0 and the second vector U1, respectively, are determined for said wall as follows: PHI0=arcsin(U0.nR) and PHI1=arcsin(U1.nR), where nR is a unit vector, normal to the wall; the distance d0 between the sound source and the microphone, as a function of the first delay TAU1, is determined by a relation of the following kind:

13. The method of claim 12, wherein d is determined by d = (TAU x C) / ((cosPHI / cosPHI)-1), where C is the speed of sound.

14. The distance z0 from the sound source to the wall is expressed by the following type of relational expression: z0=d0(sinPHI0-sinPHI1) / 2 The method of claim 13, wherein the value is determined by

15. the space comprises a plurality of walls; an inverse frequency-to-time transformation is applied to the generalized velocity vector in order to represent it in the time domain V'(t) in the form of a series of peaks (equation 39b, FIG. 2); In the series of peaks, peaks associated with a reflection at one wall of the plurality of walls are identified, and each identified peak represents a first delay TAU of a sound path caused by the reflection at a corresponding wall n relative to the straight path. n and has a time coordinate that is a function of - each first delay TAU of the first vector U0 and each second vector Un representing the sound path caused by the reflected sound at the wall n; n At least one parameter is a function of * The direction of the straight line path (DoA); *The distance d0 from the sound source to the microphone; * at least one distance zn from the sound source to the wall n.

16. 16. The method of claim 15, wherein the peaks associated with a reflection at a wall n have time coordinates that are multiples of the delay TUn associated with that wall n, and a first portion of peaks having the smallest positive time coordinates are preselected to identify each peak in that portion associated with a single reflection at a wall.

17. 17. An acoustic signal processing apparatus comprising a processing circuit for implementing the method according to any one of claims 1 to 16.

18. 17. A computer program comprising instructions which, when executed by a processor of a processing circuit, perform the method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • FR1911723

  • System of estimating direction of sound source

    JP2004012151A

  • Sound source localization based on reflections and room estimation

    US20110317522A1

  • Acquisition of spatialized sound data

    US20160277836A1

  • Location of sound sources in a given acoustic environment

    WO2019239043A1