Method and system for generating an audio scene in a binaural spatialization system

By defining left and right frontal regions and applying filter flattening to auditory transfer functions, the method enhances the subjective perception of 3D audio scenes, addressing the limitations of general auditory transfer functions in binaural spatialization systems.

WO2026109728A1PCT designated stage Publication Date: 2026-05-28MUSIC UNIT
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/083881
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2025-11-21
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing binaural spatialization systems fail to provide satisfactory subjective perception of 3D audio scenes due to the use of general auditory transfer functions for a broad audience, and methods like averaging transfer functions across individuals do not meet expectations.

Method used

A method and system that involves defining left and right frontal regions and applying filter flattening to the amplitude and/or phase of auditory transfer functions for specific ear filters, tailored to individual listeners, enhancing the subjective experience of 3D audio scenes.

Benefits of technology

The proposed method and system significantly improve the subjective perception of 3D audio scenes by providing personalized audio rendering, ensuring accurate sound localization and immersion, applicable in various audio playback configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025083881_28052026_PF_FP_ABST
    Figure EP2025083881_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating an audio scene in a binaural spatialization system, the method comprising a step of providing a set of primary auditory transfer functions, also called filters, relating to the left ear and to the right ear of individuals, for example measured on a set of individuals, the method comprising a step of defining left and right frontal regions (ZFg), and a step of flattening the filters, the amplitude and / or the phase of the left-ear filters being set to a constant or even zero value in the left frontal region, and the amplitude and / or the phase of the right-ear filters being set to a constant or even zero value in the right frontal region; the invention also relates to an associated device and associated system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and system for generating an audio scene in a binaural spatialization system

[0002]

[0001] The invention relates to a method for generating an audio scene in a binaural spatialization system. The invention also relates to an audio system and device in which the proposed method can be implemented.

[0003] In general, it is becoming increasingly common to need to generate a three-dimensional, or multi-channel, audio scene by synthesis, whether to reproduce stereophonic music listening on headphones or via earbuds or in the case of immersion scenarios in a virtual reality with a virtual reality headset or even on speakers, especially thanks to cross-path cancellation techniques.

[0004]

[0003] Binaural hearing is when both ears of an individual receive specific sound signals, corresponding to a perception of sounds in space allowing the auditory system and the brain of the individual to determine the direction from which a sound originates.

[0005]

[0004] Thus, thanks to binaural hearing, spatialized sound reproduction allows a listener to perceive sound sources coming from any direction or position in space, and to have the impression of being immersed in a realistically rendered three-dimensional sound scene.

[0006]

[0005] The signals received at the right and left ears of the listener are differentiated, in particular but not exclusively, by their intensity and by the delay of arrival and phase.

[0007]

[0006] The term "multichannel," in the context of spatialized sound reproduction, refers to producing a representation of an acoustic scene in the form of P signals (called spatial components). These signals contain all the sounds that make up the sound scene, but with weightings that depend on their direction (or "incidence") and are described by P associated spatial rendering functions. It should be noted, however, that the proposed method also works for a scene with a single audio source.

[0008]

[0007] The specific spatialized sound reproduction techniques to which the present invention relates are based on the existence of acoustic transfer functions of the head between spatial positions and the ear canal. These transfer functions, called "HRTF" (for "Head Related Transfer Functions"), concern the frequency form of the transfer functions. Their time-domain form will be referred to hereafter as "HRIR" (for "Head Related Impulse Response").

[0009] More specifically, we are particularly interested in a set of auditory transfer functions comprising for a given individual and for a plurality of measured positions M: a frequency transfer function for the left ear (HRTFg(M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRI Rg(M)(t)) and an impulse response function for the right ear (HRIRd(M)(t)).

[0010]

[0009] There are public databases in which these measurements of auditory transfer functions have been listed for a plurality of individuals. In the remainder of this document, a particular individual will be assigned the index ï and the number of individuals will by convention be equal to N.

[0011]

[0010] There are also public and private databases containing reference auditory transfer functions. The transfer functions used in this document may also have been generated by a mathematical model from previous measurement data or from theoretical data.

[0012]

[0011] The ipsilateral ear is the ear that receives the sound first, and the contralateral ear is the ear that receives the sound second for the same sound source. In the remainder of this document, the left hemisphere is the hemisphere located to the left of the sagittal plane referenced by the listener's head in an upright position, and the right hemisphere is the hemisphere located to the right of this sagittal plane.

[0013]

[0012] The left ear is the ipsilateral ear for sounds coming from the left hemisphere and the right ear is the ipsilateral ear for sounds coming from the right hemisphere.

[0014]

[0013] For the generation of a sound scene by synthesis in a binaural rendering engine, a convolution operation is carried out between the sound sources of the scene and the temporal impulse responses, respectively right and left, or if working in the frequency domain, a multiplication operation of the spectra followed by an inverse Fourier transform.

[0015]

[0014] It turns out that using a set of auditory transfer functions measured on a particular individual is not satisfactory if used for a general binaural rendering engine used for a general audience.

[0016]

[0015] Building an average of measurements over N individuals and producing a set of average auditory transfer functions does not give complete satisfaction either.

[0017]

[0016] Some have tried to work on the phase of the transfer functions as for example in document US10609504 but the subjective results still leave something to be desired.

[0018]

[0017] The inventors sought to improve the situation, in particular to improve the subjective perception of the 3D audio scene rendering.

[0019]

[0018] The proposed method and system can also be used in a so-called "transaural" context with listening on loudspeakers.

[0020]

[0019] The inventors also maintained the objective of being able to offer a possibility of individualization with respect to the listener who will be immersed in the three-dimensional audio scene.

[0021] To this end, a method is proposed here for generating an audio scene in a binaural spatialization system, the method comprising: S0 - a step of providing a set of primary auditory transfer functions, relating to the left and right ear of individual(s), the set of primary auditory transfer functions comprising for each of a plurality of measured positions Mj:

[0022] a frequency transfer function for the left ear (HRTFg(Mj)(f)) denoted Hg(9j,0j) and comprising amplitude and phase,

[0023] a frequency transfer function for the right ear (HRTFd(Mj)(f)) denoted Hd(0j, Oj) and comprising amplitude and phase,

[0024] a left-ear impulse response time function (HRI Rg(Mj)(t)) denoted h g (0j, Oj), and an impulse time response function for the right ear (HRIRd(Mj)(t)) denoted hd(0j, Oj),

[0025] each position Mj having the direction in polar coordinates azimuth

[0026]

[0027] and elevation

[0028] Frequency and time transfer functions, generally called filters, can also be the result of an averaging operation based on results from several individuals, and can be symmetrized or not.

[0029] the audio scene to be generated comprising one or more audio sources ASk, each located at a distance Rk, an azimuth 0k and an elevation Ok, k being an index that can range from 1 to Nbs (number of sources),

[0030] the process being characterized in that it comprises:

[0031] - the definition of a left frontal region (ZFg) comprising a first subset of points Mj for which -Q o <0j <0 and -Oo< 0 / <+Oo,

[0032] - the definition of a right frontal region (ZFd) comprising a second subset of points Mj for which O<0j<0o and -00<0 / <+Oo,

[0033] - a filter flattening step (SAP) in which the amplitude and / or phase of the left-ear filters are set to a substantially constant value only in the left frontal region, and the amplitude and / or phase of the right-ear filters are set to a substantially constant value only in the right frontal region,

[0034] - a usage step (SR) to generate an audio scene rendering using a binaural rendering engine based on filters with flattening and the ASk audio sources of the audio scene to be generated.

[0035]

[0021] Put another way, the left ear filters are flattened over the left frontal region of interest, that is, for a spatial portion of sounds coming from the left (therefore, the ipsilateral left ear here), but the left ear filters are not flattened for sounds coming from the right (contralateral left ear). Conversely, the right ear filters are flattened over the right frontal region of interest, that is, for a spatial portion of sounds coming from the right (ipsilateral right ear), but the right ear filters are not flattened for sounds coming from the left (contralateral right ear).

[0036]

[0022] Put another way, for a source positioned angularly in the left frontal region of interest, the left (ipsilateral) ear filter undergoes the proposed flattening, while the right (contralateral) ear filter remains unchanged. Conversely, for a source positioned angularly in the right frontal region, the right (ipsilateral) ear filter undergoes the proposed flattening, while the left (contralateral) ear filter remains unchanged.

[0037]

[0023] It is therefore understood that the flattening is asymmetrical. In practice, this gives surprisingly satisfactory results.

[0038]

[0024] The flattening step may only involve flattening the filter's amplitude without affecting the phase portion. In an alternative embodiment, the flattening step involves both the amplitude and the phase of the filter.

[0039]

[0025] It should be noted that during the usage stage (SR), the generation of the audio scene rendering can be done in real time, or conversely, this audio scene rendering can be stored and played back later. In this document, azimuth angles 0 are expressed with positive values ​​on the right side and negative values ​​on the left side, 0 = 0 being the straight-ahead direction.

[0040] In practice, filters can be in the so-called "sofa" format. The constant value used in the flattening step applies to all frequencies of the audio spectrum (either for amplitude only or for amplitude and phase as discussed previously).

[0041]

[0029] In a particular case, this value is zero, in other words, a complete selective bypass of the filter is carried out in the front, left and right regions.

[0042] Note that 0o and Oo are predefined parameters.

[0043] Regarding the first and second subsets of points in the frontal, left and right regions: these could be all points Mj for which -00 < 0j < 0 (resp. 0 < 0j < 00) and -00 < 0 / < +00. Alternatively, they could be points with a direction circumscribed within a circular or elliptical cap shape.

[0044]

[0032] As will be seen later, the distance of the ASk source can also condition the application of the flattening step. In other words, the front left and right regions can be defined independently of the distance of the source composing the audio scene to be reproduced, or alternatively, the front left and right regions can have a certain limited depth, as will be seen below, forming angular sectors finite in depth.

[0045]

[0033] It is noted that the calculations can be done in linear or logarithmic mode using decibels in dB.

[0046]

[0034] According to one embodiment, the flattening step (SAP) is performed in real time. An "on-line" switching is performed, that is to say, the filter is bypassed for sources that are located inside the front flattening regions.

[0047]

[0035] According to one embodiment, the flattening step (SAP) is performed in offline mode (in practice, this is referred to as 'OFFLINE' mode). The filter values ​​in the memory storage area are replaced with those for the directions located within the frontal flattening regions. The filters thus modified are then used for the real-time binaural engine.

[0048] According to one implementation, the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

[0049]

[0037] For example, in practice, the flattening step is carried out for distances less than a predetermined threshold radius Rma, and the flattening step is not carried out for distances greater than this predetermined threshold Rma.

[0050] According to one design, the process involves smoothed joints at the edge of the front regions, respectively left and right.

[0051] The smoothing in question can be done in time (temporal smoothing) or in space (transition band).

[0052] In the configuration where flattening is applied only to the amplitude, this solution proves effective; in other words, a "cross-fade" process is used. Combing effects are avoided at the edges of the flattened front regions. The blending may involve, on a boundary strip, an interpolation between the filter values ​​just outside the flattened area and the flattening constant.

[0053]

[0041] According to one embodiment, a further step, denoted SD, may be provided for evaluating the interaural delay, denoted T(0, ), the average interaural delay T(9, CI)) is calculated as follows:

[0054] I_l(θ,φ) = (1 / N)∑ᵢᴺ argmax(|h_g^i(θ,φ)*w|)

[0055] I_r(θ,φ) = (1 / N)∑ᵢᴺ argmax(|h_d^i(θ,φ)*w|)

[0056] 1 /

[0057] τ(θ,φ) = (1 / fs)(I_r(θ,φ) - I_l(θ,φ))

[0058]

[0059] fs being the sampling frequency of the HRIRs, and argmax^) being the index of the maximum of a sequence u n * being the convolution product, w being a convolution filtering function,

[0060] with time-domain application on the impulse response or application in the frequency domain, the filter is then calculated as follows:

[0061] HH_l(θ,φ) =

[0062]

[0063]

[0042] Correction by intraural delay is particularly relevant in the case of flattening only on the amplitude.

[0064]

[0043] According to one embodiment, it may further provide for a step of individualizing the intraural delay (SG) as a function of a head radius b of a target listener, the intraural delay being determined by - aT, with b the new head radius and a the original radius.

[0065]

[0044] According to one embodiment, an angular correction step (SH) may be further provided, with a parameter y that can be adapted to each user, according to the formulation:

[0066] f(0, <>) = (0 + y * sin(20), 0)

[0067]

[0045] According to one embodiment, a further step, denoted SB, may be provided, for smoothing the spectral amplitude by binaural relative deviation, and normalization for the reference direction (0 O , O ) with an average across the subjects, according to the following formulation:

[0068] N

[0069] dB(H_l(θ,φ)) = (1 / N)∑ᵢ(L^i(θ,φ) - L^i(θ₀,φ₀))

[0070] i

[0071] N

[0072] dB(H_r(θ,φ)) = (1 / N)∑ᵢ(R^i(θ,φ) - L^i(θ,φ)) + dB(H_l(θ,φ))

[0073]

[0074] i

[0075]

[0046] The invention also relates to an audio apparatus or device configured to carry out the steps of a process as described above. Depending on the configuration, the audio apparatus or device may (or may not) perform the flattening step.

[0076]

[0047] A head tracking function may be provided that supplies, in real time, a signal denoted HDT(t) representing the head movement in yaw rotation, as well as in left-right lateral tilt and front-back tilt, thus three spatial rotation coordinates. The relative positions of the scene sources are corrected by compensating for the detected movements and displacements.

[0077] The invention also relates to a processing system configured to implement the steps of a process as described above.

[0049] The invention will be further detailed by describing non-limiting embodiments, and based on the accompanying figures illustrating variants of the invention, in which:

[0078] - [Fig.1] is a schematic diagram illustrating the spherical coordinates of points and directions in space around an individual, for example a listener in the context of listening or in the context of measuring auditory HRTF functions;

[0079] - [Fig.2] is analogous to figure 1 in the case of an audio scene rendering and illustrates in particular an example of a right angular front region where the filters are flattened for the right ear filters;

[0080] - [Fig.3] is analogous to figure 2 and illustrates in particular an example of left angular frontal region where the filters are flattened for left ear filters;

[0081] - [Fig.4] is analogous to figure 3 and illustrates a particular example of left angular frontal region where the filters are flattened for left ear filters, the frontal region having a capped / limited depth;

[0082] - [Fig.5] schematically illustrates an audio scene reproduction using a binaural rendering engine, for a single ASi audio source;

[0083] - [Fig.6] illustrates an example of a frequency transfer function, for the right ear, with the amplitude and phase for four different azimuth positions and for a given elevation position;

[0084] - [Fig.7] illustrates the example of the transfer function of figure 6 having undergone the flattening step only in amplitude;

[0085] - [Fig.8] illustrates the example of the transfer function in Figure 6 having undergone the amplitude and phase flattening step;

[0086] - [Fig.9] schematically illustrates an example of the sequence of steps of the proposed process according to a first variant of hardware implementation;

[0087] - [Fig.10] schematically illustrates an example of the sequence of steps of the proposed process according to a second variant of hardware implementation;

[0088] - [Fig.11] schematically illustrates a more complete example of the sequence of steps of the proposed process according to a third variant.

[0089]

[0050] In the various figures, the same reference numerals designate identical or similar elements. For the sake of clarity, some elements are not necessarily shown to scale.

[0090]

[0051] Generalities and system

[0091]

[0052] Generally, spatialized binaural sound reproduction allows a listener U to be immersed in an audio scene and to perceive sound sources coming from different directions and positions in space, while the listener U only has two speakers HPg and HPd.

[0092]

[0053] In the example of Figure 5, the playback device is an audio headset 2 with the left speaker HPg arranged opposite the left ear OG of the listener and the right speaker HPd arranged opposite the right ear OD of the listener U.

[0093]

[0054] In other configurations, the rendering organ can be a pair of intra-aural earbuds such as, for example, “earbuds”.

[0094]

[0055] In yet another configuration, in a so-called open-air listening context sometimes referred to as 'transaural™', two reproduction devices, such as conventional loudspeakers, are arranged with crosstalk cancellation. The proposed invention can also be applied in this transaural configuration, with two or more loudspeakers.

[0095] For the different configurations mentioned above, both speakers can be controlled by a single control unit directly in wired mode, or the speakers can be controlled by a single control unit but via a wireless Bluetooth connection.

[0096]

[0057] In one scenario, the audio scene to be reproduced may consist of a single sound source, denoted ASi in Figure 5. In another scenario, the audio scene to be reproduced consists of a plurality of sound sources ASi. For example, the audio scene to be reproduced may include NbS sound sources ASk (k ranging from 1 to N).

[0097]

[0058] Each sound source to be reproduced, such as the first source ASi, is located at a position M1 characterized by its polar coordinates O1 and < D1 and its distance R1 (also denoted r1), as illustrated in Figure 1. Polar coordinates will be used throughout this document. The first coordinate of a generic point M O represents the azimuth with respect to a "straight ahead" direction, with positive angles taken to the right. The second coordinate O (which can also be denoted tp) corresponds to the elevation, also called "site," with respect to a horizontal direction, with positive angles taken upwards.

[0098]

[0059] In Figure 1, the plane labeled PH is the horizontal reference plane for the listener with their head upright. The plane labeled PMS is the median sagittal plane, separating the left hemisphere from the right hemisphere. The plane labeled PVI is the vertical internal plane (passing through 0° = 90°). When the system is equipped with the head-tracking function, the deviation of the actual position from the upright head position is compensated by the same angle values, but with opposite signs. For the listener, the sound source always appears to be in its given position in space, regardless of any small head movements the listener may make.

[0099] The reproduction of an audio scene is based on the use of acoustic transfer functions of the head between the positions in space (0,0) and the ear canal, called auditory transfer functions. These transfer functions, known as "HRTFs" (for "Head-Related Transfer Functions"), relate to the frequency form of the transfer functions. Their time-domain form, that is, the impulse response, will be referred to hereafter as "HRIRs" (for "Head-Related Impulse Response").

[0100]

[0061] In the context of binaural playback, the following transfer functions are then used: a frequency transfer function for the left ear (HRTFg(M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRI Rg(M)(t)) and an impulse response function for the right ear (HRI Rd(M)(t)), M being the position of the source in polar coordinates, f the frequency and t the time.

[0101] Transfer functions are also referred to as 'filters' in this document.

[0102] The filters used in the process described below may or may not be symmetrical; they may have been averaged over N individuals, or they may originate from a single individual. These filters may be public or proprietary.

[0064] In this document, the frequency transfer functions are manipulated in digitized / numeric form, namely as a sequence of values. For a large number of frequencies within the human audible spectrum, the transfer function provides an amplitude and a phase. Typically, 128 frequency points may be listed, or even 256. For each point, the amplitude and phase are encoded in 8, 10, or 12-bit digital format.

[0103] Transfer functions can be coded in the 'sofa' format known in the art of sound engineers.

[0104] Returning to Figure 5, the audio scene is reproduced by a binaural rendering engine. This engine has access to transfer functions processed in either their frequency or time domain form. Using these functions, for each sound source ASi in the audio scene to be reproduced—characterized by the time-wave pattern ASi(t) of the sound source, its polar coordinates (0i, Oi), and its distance from the listener Ri—the binaural rendering engine generates the corresponding three-dimensional audio scene in the two loudspeakers.

[0105]

[0067] Figure 5 represents only one sound source (01, O1, R1) but the reader will understand that what follows will be repeated for each of the sound sources, and that the respective signals of the different sound sources are added together in the waves produced towards the loudspeakers (HPg, HPd).

[0106] When the binaural rendering engine works in the frequency domain, from the time pattern ASi(t) a Fast Fourier Transform (FFT) is calculated, denoted XASi(f), more simply denoted XAS in figure 5. The coordinates (01, O1,r1) of the sound source are applied on one side to the HHi filter for the left ear and on the other side to the HHr filter for the right ear.

[0107] HHi and HHr are the filters resulting from upstream processing, with flattening where applicable, or flattening being done directly by the binaural rendering engine.

[0108]

[0070] For the left channel, the fast Fourier transform XAS of the source signal is multiplied by the filter HHi, which is assigned the coordinates of the source. This gives the signal to be reproduced in frequency form, WASg. This signal is then subjected to an inverse Fourier transform calculation 21, thus obtaining the time-domain signal WASg(t), i.e., the image of the sound wave, which will be played on the left loudspeaker HPg.

[0109]

[0071] For the right channel, a similar procedure is used: the fast Fourier transform of the source signal is multiplied by the filter HHr, assigned the coordinates of the source, which gives the signal to be reproduced in frequency form WASd. This signal is then subjected to an inverse Fourier transform calculation 22, thus obtaining the time-domain signal WASd(t), i.e., the image of the sound wave, which will be played on the right loudspeaker HPd.

[0110]

[0072] In an alternative implementation, when the binaural rendering engine works in the time domain, for the left channel a convolution product is performed between the sound source signal ASi(t) and the impulse response function hhi(t) with coordinates (01,01,r1) of the sound source, thus obtaining the time signal WASg(t) which will be played on the left speaker HPg.

[0111]

[0073] A similar procedure is used for the right channel; a convolution product is performed between the sound source signal ASi(t) and the impulse response function hh r (t) affected by the coordinates (01, O1,r1) of the sound source we thus obtain the time signal WASd(t) which will be played on the right speaker HPd.

[0112]

[0074] It should be noted that to a certain extent the left and right transfer functions (HHi and HHr) can be customized for the listener experiencing the three-dimensional audio scene. This includes adapting the head diameter to internal delays, azimuthal angular correction ('gamma' correction), or combined azimuthal and elevation angular correction.

[0113]

[0075] Finally, regarding Figure 5, if the head tracking function is present, it provides in real time a signal denoted HDT(t) representing the head movement in yaw rotation, as well as in left-right lateral tilt and front-back tilt, thus three spatial rotation coordinates. As explained elsewhere, the rendering engine compensates by injecting these values ​​with opposite signs into the sound signal generation algorithm.

[0114]

[0076] Advantageously, according to the present invention, the filters are flattened or flattened. This is illustrated in particular in Figure 2 and Figure 3.

[0115]

[0077] A right frontal region ZFd is defined, visible in Figure 2, in which the amplitude and / or phase of the right ear filters are set to a constant value.

[0116]

[0078] Furthermore, a left frontal region ZFg is defined, visible in Figure 3, in which the amplitude and / or phase of the left ear filters are set to a constant value.

[0117]

[0079] The left frontal region ZFg encompasses a first subset of points Mj for which -0° < 0j < 0 and -0° < 0j < +0°. The right frontal region ZFd encompasses a second subset of points Mj for which 0 < 0j < 0° and -0° < 0j < +0°. It is recalled here that the azimuth angles 0° are expressed with positive values ​​on the right side and negative values ​​on the left side, 0° = 0 being the straight-ahead direction.

[0118] However, the inverse reference for the sign of 0 could be used, mutatis mutandis.

[0119] The step of defining the frontal regions ZFg and ZFd is marked DRF in figures 9 to 11.

[0120] The flattening performed may only affect the amplitude of the filter without affecting the phase portion. In an alternative embodiment, the flattening step affects both the amplitude and the phase of the filter.

[0121]

[0083] Regarding the first and second subsets of points in the frontal, left and right regions, it can be all the points Mj for which -0o <0j <0 (resp. O<0j<0o ) and -0o<0j<+0o, in which case we have a rectangular region as illustrated in the figures.

[0122]

[0084] Alternatively, the first and second subsets of points may concern points with a direction circumscribed in a circular or elliptical cap shape, as illustrated by the dotted line ZF'.

[0123]

[0085] The constant value used in the flattening step applies to all frequencies of the audio spectrum (either for amplitude only or for amplitude and phase as discussed previously).

[0124] In a particular case, this value is zero; in other words, a complete selective bypass of the filter is performed in the front, left, and right regions.

[0087] The reference angles 0o and Oo are configurable. The reference angles 0o and Oo can correspond to 0o = 15° and Oo = 15°. These are referred to as 'image' filters in the context of the present invention. According to another solution, the reference angles 0o and Oo can correspond to 0o = 30° and Oo = 15°. These are referred to as 'audio' filters in the context of the present invention.

[0125] According to a basic embodiment, filter flattening is applied to the aforementioned front regions without taking into account the distance from the audio source to be reproduced.

[0126]

[0089] However, according to one option of the method, the distance of each audio source can play a role in the filter flattening process. In other words, the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

[0127]

[0090] In practice, referring to Figure 4 which illustrates this configuration, the flattening step is carried out for distances less than a predetermined threshold radius Rma, and the flattening step is not carried out for distances greater than this predetermined threshold Rma.

[0128]

[0091] As an alternative to a binary logic with respect to the distance from the source, it is not excluded to use a progressive modulation, e.g. we flatten less and less as the distance increases.

[0129] In figures 6 to 8, which illustrate filters, the x-axis represents frequency, graduated logarithmically. The y-axis represents amplitude and phase in decibels, respectively.

[0130] Figure 6 illustrates an example of a frequency transfer function, for the right ear, with the amplitude and phase for four given azimuth positions and for one azimuth position, without a flattening step.

[0131]

[0094] Curve 81 (solid line) represents the amplitude of the frequency transfer function, for the right ear, for 0 = 0°.

[0132] Curve 82 (small dotted line) represents the amplitude of the frequency transfer function for 0 = 15°. Curve 83 (dashed line) represents the amplitude of the frequency transfer function for 0 = 30°. Curve 84 (solid line) represents the amplitude of the frequency transfer function for 0 = 45°.

[0133] Curve 85 represents the phase of the frequency transfer function for the right ear at 0°. Curve 86 (small dotted line) represents the phase of the frequency transfer function at 15°. Curve 87 (dashed line) represents the phase of the frequency transfer function at 30°. Curve 88 represents the phase of the frequency transfer function at 45°.

[0134] We have taken Φ=0° as the reference for the filter examples here.

[0135]

[0136] Figure 7 reproduces the transfer function example from Figure 6, applying the flattening step only in amplitude. According to the illustrated example, the reference angle is 0o = 35°.

[0137]

[0099] For the amplitude, curves 81 to 83 are replaced by a flat curve 91, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. On the other hand, curve 84 for 0 = 45° is not modified and remains unchanged (direction outside ZFd).

[0138]

[0100] The phase curves remain unchanged.

[0139] Figure 8 reproduces the example of the transfer function from Figure 6 with application of the flattening step in amplitude and phase.

[0102] For amplitude, curves 81 to 83 are replaced by a flat curve 91, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. However, curve 84 for 0 = 45° is not modified and remains unchanged.

[0140]

[0103] For the phase, curves 85 to 87 are replaced by a flat curve 92, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. On the other hand, curve 88 for 0 = 45° is not modified and remains unchanged.

[0141]

[0104] If for the same filters we took as reference angle 0o = 22°, the curves 81 to 82 would be replaced by the flat curve 91, but the curves 83 and 84 would remain unchanged (directions outside ZFd).

[0142]

[0105] The method advantageously provides for smoothed connections at the location of the border of the front regions ZFg and ZFd, respectively left and right.

[0143]

[0106] The smoothing in question can be carried out in time (temporal smoothing) or in space (transition band).

[0144]

[0107] The different stages of reprocessing these transfer functions are illustrated in figures 9 to 11. The first stage, noted S0, corresponds to making available the above-mentioned transfer functions from databases or measurements with, where appropriate, right-left symmetrization, and where appropriate, averaging over N individuals.

[0145]

[0108] In the generic example of Figures 9 and 10, the step of defining the DRF flattening front areas is followed by the step of flattening the SAP filters. In this example, there is no other correction, and the result of the flattening is used to run the binaural rendering engine (SR step).

[0146]

[0109] According to the first possibility, illustrated in Figure 10, the SAP flattening step is performed in real time. An "on-line" switching is performed, that is to say, the filter is bypassed for sources that are located inside the front flattening regions.

[0147] According to the second possibility, illustrated in Figure 9, the SAP flattening step is performed offline. The filter values ​​for directions located within the frontal flattening regions are then replaced in a memory storage area. These modified filters are then used by the real-time binaural engine.

[0148] Furthermore, a spectral amplitude symmetrization step SA can be planned, including additional smoothing by local average of k neighboring directions j, according to the formulation:

[0149] I

[0150] W = IZ / \ dB + dB (w i Hid ( ~ 9 I

[0151] ■ ^p)

[0152] I

[0153] W

[0154]

[0155] = |Z j \ dB ( WjHid ( 9j ' + dB 0 p) According to one option, the process provides for smoothing the spectral amplitude by binaural relative deviation, normalization for a reference direction (0 O , <>0) with average across subjects, according to the formulation, for a source on the left:

[0156] N

[0157] dB (H t (0, 0)) = ( / / (0, 0) - L l (0 O , 0 O ))

[0158]

[0159] N dBÇHAe.iï) = (^(0,0) - r (0,0)) + dB(w z (0,0)

[0160]

[0161] i

[0162] This step is denoted SB. For a right-hand source, normalization with respect to the reference direction 0o, Oo right-hand side, can be applied similarly:

[0163] N

[0164] dB (H r (0, 0)) = — R 1 (00, 0 O ))

[0165] i

[0166] N

[0167] dBÇHtCe.iï) = (r(0,0) - R^B, 0)) + dB(H r (0,0)

[0168] i

[0169]

[0170] We note that the flattening of the filters on the frontal regions ZFg ZFd can also be expressed mathematically according to the following formulation:

[0171] N

[0172] dBÇHtCB.iï) = - ^ (r (0,0) - Z / (|0|, 0))

[0173] i N

[0174] dB(H r (9, )') = -^ (R ; (0,0) - r(0,0)) + dB(W z (0,0)

[0175]

[0176] i

[0177] Furthermore, a step denoted SC can be performed for symmetrization and phase averaging, according to the following formulation:

[0178] N

[0179] < H^e, & > = (< Wj(0, 0) > + < H d ' ((- 0., 0.) > )

[0180] i

[0181] N

[0182] < H r (e, 0) > = ^ (< W' (0, 0) > + < H g ((- 0.0.) > )

[0183]

[0184] i

[0185] The notation <. > denotes the phase of the complex transfer function. Furthermore, a step denoted SD can be carried out, evaluating the average intraaural delay denoted T(0,0), the average intraaural delay T(9, CI)) being calculated as follows;

[0186] 1 N

[0187] Ii(9,c / >') = ~^ argmax(]h g l (9, ( / )') * w|)

[0188] i

[0189] N

[0190] L r (0,(l)) = - 2^ argmax(\h d l(9, <l)) * w|)

[0191] i

[0192] T(0,0) = — ( / r (0,0) - / z (0,0)

[0193]

[0194] fs being the sampling frequency of the HRIR impulse responses, and argmax(u n ) being the index of the maximum of a sequence u n , digitized values ​​of the impulse response, the notation * being the convolution product, w being a convolution filtering function. The convolution filtering function w can be reduced to a Dirac pulse (then we simply select the largest signal peak).

[0195] The mean intraaural delay can be applied in the time domain, by applying a corresponding delay to the impulse response.

[0196]

[0120] The average interaural delay can be applied in the frequency domain, the filter then being calculated as follows: HH l (6, <p') = 10 20

[0197] [

[0198]

[0199] 122] HH r (0,) = io dB ( H r( 0 ^)) / 2O e ZT ^^)

[0200]

[0123] Furthermore, an individualization step of the interaural delay, denoted SG, can be carried out as a function of a head radius b of a target listener, the interaural delay being determined by -T, with b the new head radius and aa

[0201] the original section.

[0202] The radius may have originated from measurements accompanying the transfer function measurements. The radius may have originated from anthropological data.

[0203] The radius b can be chosen according to the listener benefiting from the audio scene rendering. For example, the radius b can be smaller than a for a listener with a small head circumference; we would refer to size XS or S.

[0204]

[0126] For example, radius b may be larger than radius a for a listener with a large head circumference, we will speak of size L or XL.

[0205] For example, radius b may be larger than radius a for a listener with an average head circumference.

[0206]

[0128] Then we proceed to a step denoted SH of angular correction.

[0207]

[0129] According to a first simple example referred to here as 'in two dimensions', an azimuthal angular correction is performed, with a parameter y that can be adapted to each user, according to the formulation:

[0208] f

[0209]

[0210] (9, 0) = (0 + y * sin(20),

[0211]

[0130] In other words, instead of using the azimuth angle 9 directly, we use instead a transform of the azimuth angle in the form 9 + y * sin(20).

[0212] The value of y is configurable.

[0213]

[0132] If the 'head movement tracking' function is implemented, the SH step performs yaw movement corrections on 0 in a manner similar to that shown above.

[0214]

[0133] According to another example of the SH angular correction step, azimuth and elevation are corrected according to a so-called 'three-dimensional' deformation, using f(0,<>) = (0 + y * sin(20),<> +y' * sin(<>)). The parameters y and y' can be adapted to each user.

[0215]

[0134] In an even more generic way, the step denoted SH of angular correction consists of replacing the angles 0, <$> in the HRTF formulas by the angles ( 0 ', <>') such that ( 6 ', <>') = f ( 6, <>) f being any azimuth and elevation correction function.

[0216]

[0135] Figure 11 shows a diagram of the process with the aforementioned options. The input metadata, namely the initial transfer filters / functions, and the definition of the right and left front ends (steps S0 and DRF), are shown at the beginning. The subsequent steps, namely SB, SAP, SD, SF, and SG, are not necessarily applied or performed in the order presented.

[0217]

[0136] Other considerations

[0218]

[0137] Optionally, due to the symmetrization, it is possible to keep the filters in memory only for one hemisphere, right or left, and to use a planar symmetry with respect to the PMS sagittal plane performed by the real-time binaural rendering engine.

[0219]

[0138] More precisely, in practice, after the phase and amplitude symmetrization steps SA and SC, data can be stored for only one hemisphere, and the complementary transformations performed only on that hemisphere. The process for the other hemisphere will be reconstructed in real time by planar symmetry with respect to the sagittal plane PMS. In this way, the memory space occupied can be halved, which may be relevant for certain categories of audio hardware or processing chains.

[0220]

[0139] According to another optional feature, in addition to the transformations described above, a further transformation may be provided consisting of collectively boosting high-pitched sounds across all filters. Thus, for example, sounds with frequencies above 2 or 3 kilohertz may be boosted, with this boost having an amplitude between 1 dB and 3 dB, for example.

Claims

DEMANDS 1. A method for generating an audio scene in a binaural spatialization system, the method comprising: S0- a step of making available a set of primary auditory transfer functions, relating to the left and right ear of individual(s), the set of primary auditory transfer functions comprising for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(Mj)(f)) denoted Hg(0j, Oj) and comprising amplitude and phase, a frequency transfer function for the right ear (HRTFd(Mj)(f)) denoted Hd(0j, Oj) and comprising amplitude and phase, a time-pulse response function for the left ear (HRIRg(Mj)(t)) denoted h g (0j, Oj), and an impulse time response function for the right ear (HRIRd(Mj)(t)) denoted h d (0j, Oj), each position Mj having the direction in polar coordinates azimuth 0j and elevation O y , The frequency and time transfer functions, generally called filters, which may also be the result of an averaging operation from results of several individuals, and which may or may not be symmetrical, the audio scene to be generated comprising one or more audio sources ASk, each located at a distance Rk, an azimuth 0k and an elevation Ok, the process being characterized in that it comprises: - the definition of a left frontal region (ZFg) comprising a first subset of points Mj for which -0 O <0j <0 and -Oo< 0 / <+Oo, - the definition of a right frontal region (ZFd) comprising a second subset of points Mj for which O<0j<0o and -00<0 / <+Oo, - a filter flattening step (SAP) in which the amplitude and / or phase of the left ear filters are set to a constant value only on the left frontal region, and the amplitude and / or phase of the right ear filters are set to a constant value only on the right frontal region, - a usage step (SR) to generate an audio scene rendering using a binaural rendering engine based on the flattened filters and the audio sources of the audio scene to be generated.

2. A method according to claim 1, wherein the flattening step (SAP) is carried out in real time.

3. Method according to claim 1, wherein the planing step (SAP) is carried out in offline mode.

4. A method according to any one of claims 1 to 3, wherein the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

5. A method according to any one of claims 1 to 4, wherein smoothed joints are provided at the edge of the front regions, left and right respectively.

6. A method according to any one of claims 1 to 5, wherein a step denoted SD is provided for evaluating the average internal delay denoted T(0, ), the average internal delay T(9, CI) is calculated as follows: TV = -^2_ l ar9max ^ h 3 ( ^ e '^ * w L) i N irte.fà = - 2_ l ar 9 max () h d(. e ' ( L ) ) * w L) i 1 / T(0, 0) = — ( / r (e, 0) - A(e, 0)) fs being the sampling frequency of the HRIRs, and argmax^) being the index of the maximum of a sequence u n, * being the convolution product, w being a convolution filtering function with time application on the impulse response or application in the frequency domain, the filter is then calculated as follows: HH_l(θ,φ) = 7. A method according to any one of claims 1 to 6, further comprising a step of individualizing the intraural delay (SG) as a function of a head radius b of a target listener, the intraural delay being determined by with b the new head radius and a the original radius from the measurements accompanying the transfer function measurements H g (0j, Oj) and Hd(0j, Oj).

8. A method according to any one of claims 1 to 7, further comprising an angular correction step, denoted SH, with a parameter y that can be adapted to each user, according to the formulation: (Q) <p') = f 6, <p~) = (0 + y * sin(20), ) the angular correction consisting of replacing the angles 0, <;> in the HRTF formulas with the angles ( 0 ( / >').

9. Audio device configured to implement the steps of a method according to any one of claims 1 to 8.

10. Processing system configured to implement the steps of a process according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Binaural synthesis, head-related transfer functions, and uses thereof

    EP0746960B1

  • Apparatus and method for head-related transfer function compression

    EP4231668A1

  • Audio signal processing method and apparatus for binaural rendering using phase response characteristics

    US10609504B2