Method and system for generating an audio scene in a binaural spatialization system

By defining left and right frontal regions with constant filter values, the method improves the subjective perception and individualization of three-dimensional audio scenes in binaural spatialization systems, addressing the limitations of existing technologies.

FR3169046A1Pending Publication Date: 2026-05-29MUSIC UNIT

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
MUSIC UNIT
Filing Date
2024-11-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for generating three-dimensional audio scenes using binaural spatialization systems fail to provide satisfactory subjective perception and individualization for a general audience, as using a set of auditory transfer functions measured on a particular individual or averaging over multiple individuals does not meet the desired quality standards.

Method used

A method involving the definition of left and right frontal regions, where the amplitude and/or phase of ear filters are set to a constant value within these regions, and a binaural rendering engine is used to generate the audio scene, allowing for individualization and improved subjective perception.

Benefits of technology

The proposed method enhances the subjective perception of three-dimensional audio scenes by providing individualized audio rendering, ensuring accurate sound localization and immersion, applicable to various audio playback configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for generating an audio scene in a binaural spatialization system, the method comprising a step of providing a set of primary auditory transfer functions, also called filters, relating to the left and right ears of individuals, for example, recorded from a group of individuals, the method comprising a step of defining left and right frontal regions (ZFg), and a step of flattening the filters, the amplitude and / or phase of the left-ear filters being set to a constant or even zero value on the left frontal region, and the amplitude and / or phase of the right-ear filters being set to a constant or even zero value on the right region, associated device and associated system. Abstract figure: Fig. 4
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and system for generating an audio scene in a binaural spatialization system

[0001] The invention relates to a method for generating an audio scene in a binaural spatialization system. The invention also relates to an audio system and device in which the proposed method can be implemented.

[0002] In general, it is increasingly common to need to generate a three-dimensional, or multi-channel, audio scene by synthesis, whether to reproduce stereophonic music listening on headphones or via earbuds or in the case of immersion scenarios in virtual reality with a virtual reality headset or even on speakers, in particular thanks to cross-path cancellation techniques.

[0003] Binaural hearing is when both ears of an individual receive specific sound signals, corresponding to a perception of sounds in space allowing the auditory system and the brain of the individual to determine the direction from which a sound originates.

[0004] Thus, thanks to binaural hearing, spatialized sound reproduction allows a listener to perceive sound sources coming from any direction or position in space, and to have the impression of being immersed in a realistically rendered three-dimensional sound scene.

[0005] The signals received at the right and left ears of the listener are differentiated, in particular but not exclusively, by their intensity and by the delay of arrival and phase.

[0006] The term "multichannel," in the context of spatialized sound reproduction, refers to producing a representation of an acoustic scene in the form of P signals (called spatial components). These signals contain all the sounds that make up the sound scene, but with weightings that depend on their direction (or "incidence") and are described by P associated spatial rendering functions. It should be noted, however, that the proposed method also works for a scene with a single audio source.

[0007] The specific spatialized sound reproduction techniques to which the present invention relates are based on the existence of acoustic transfer functions of the head between spatial positions and the ear canal. These transfer functions, called "HRTFs" (for "Head Related Transfer Functions"), concern the frequency form of the transfer functions. Their time-domain form will be referred to hereafter as "HRIRs" (for "Head Related Impulse Response").

[0008] More particularly, we are particularly interested in a set of auditory transfer functions comprising for a given individual and for a plurality of measured positions M: a frequency transfer function for the left ear (HRTFg (M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRIRg(M)(t)) and an impulse response function for the right ear (HRIRd(M)(t)).

[0009] Public databases exist in which these measurements of auditory transfer functions have been listed for a plurality of individuals. In the remainder of this document, a particular individual will be assigned the index 'i' and the number of individuals will by convention be equal to N.

[0010] There are also public and private databases containing reference auditory transfer functions. The transfer functions used in this document may also have been generated by a mathematical model from previous measurement data or from theoretical data.

[0011] The ipsilateral ear is the ear that receives the sound first, and the contralateral ear is the ear that receives the sound second for the same sound source. In the remainder of this document, the left hemisphere is the hemisphere located to the left of the sagittal plane referenced by the listener's head in an upright position, and the right hemisphere is the hemisphere located to the right of this sagittal plane.

[0012] The left ear is the ipsilateral ear for sounds coming from the left hemisphere and the right ear is the ipsilateral ear for sounds coming from the right hemisphere.

[0013] For the generation of a sound scene by synthesis in a binaural rendering engine, a convolution operation is carried out between the sound sources of the scene and the temporal impulse responses, respectively right and left, or if working in the frequency domain, a multiplication operation of the spectra followed by an inverse Fourier transform.

[0014] It turns out that using a set of auditory transfer functions measured on a particular individual is not satisfactory if used for a general binaural rendering engine used for a general audience.

[0015] Building an average of measurements over N individuals and producing a set of average auditory transfer functions does not give complete satisfaction either.

[0016] Some have tried to work on the phase of the transfer functions as for example in document US 10609504 but the subjective results still leave something to be desired.

[0017] The inventors sought to improve the situation, in particular to improve the subjective perception of the 3D audio scene rendering.

[0018] The proposed method and system can also be used in a so-called "transaural" context with listening on loudspeakers.

[0019] The inventors also maintained the objective of being able to offer a possibility of individualization with respect to the listener who will be immersed in the three-dimensional audio scene.

[0020] To this end, a method for generating an audio scene in a binaural spatialization system is proposed herein, the method comprising: S0 - a step of making available a set of primary auditory transfer functions, relating to the left and right ear of individual(s), the set of primary auditory transfer functions comprising for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(Mj)(f)) denoted Hg (0j, <e>j) and including amplitude and phase, a frequency transfer function for the right ear (HRTFd(Mj)(f)) denoted Hd(0j , <e>j) and including amplitude and phase, a time impulse response function for the left ear (HRIRg(Mj) (t)) denoted hg(0j, <e>j), and an impulse time response function for the right ear (HRIRd(Mj) (t)) denoted hd(0j,O>j), each position Mj having the direction in polar coordinates azimuth Qj and elevation Frequency and time transfer functions, generally called filters, can also be the result of an averaging operation based on results from several individuals, and can be symmetrized or not. the audio scene to be generated comprising one or more audio sources ASk, each located at a distance Rk, an azimuth 0k and an elevation <e>k, where k is an index that can range from 1 to Nbs (number of sources), the process being characterized in that it comprises: - the definition of a left frontal region (ZFg) comprising a first subset of points Mj for which -00 <07 <0 and - <e> 0 < ¢ / <+ <e>0 - the definition of a right frontal region (ZFd) comprising a second subset of points Mj for which O<07<00 and - <e>0 -cC^+Oo, - a filter flattening step (SAP) in which the amplitude and / or phase of the left-ear filters are set to a constant value over the left frontal region, and the amplitude and / or phase of the right-ear filters are set to a constant value over the right frontal region, - a usage step (SR) to generate an audio scene rendering using a binaural rendering engine based on filters with flattening and the ASk audio sources of the audio scene to be generated.

[0021] Put another way, the left ear filters are flattened over the left frontal region of interest, that is, for a spatial portion of sounds coming from the left (therefore, the ipsilateral left ear here), but the left ear filters are not flattened for sounds coming from the right (contralateral left ear). Conversely, the right ear filters are flattened over the right frontal region of interest, that is, for a spatial portion of sounds coming from the right (ipsilateral right ear), but the right ear filters are not flattened for sounds coming from the left (contralateral right ear).

[0022] Put another way, for a source positioned angularly in the left frontal region of interest, the left (ipsilateral) ear filter undergoes the proposed flattening, while the right (contralateral) ear filter remains unchanged. Conversely, for a source positioned angularly in the right frontal region, the right (ipsilateral) ear filter undergoes the proposed flattening, while the left (contralateral) ear filter remains unchanged.

[0023] The flattening step may only involve flattening the filter's amplitude without affecting the phase portion. In an alternative embodiment, the flattening step involves both the amplitude and the phase of the filter.

[0024] It should be noted that at the usage stage (SR) the generation of the audio scene rendering can be done in real time, or conversely this audio scene rendering can be stored and can be played later.

[0025] In this document, azimuth angles 0 are expressed with positive values ​​on the right side, and negative values ​​on the left side, 0 = 0 being the straight ahead direction.

[0026] In practice in the trade, the filters can be in the so-called “sofa” format.

[0027] The constant value used in the flattening step applies to all frequencies of the audio spectrum (either for amplitude only or for amplitude and phase as discussed previously).

[0028] In a particular case, this value is zero, in other words, a complete selective bypass of the filter is carried out in the front, left and right regions.

[0029] We note that 00 and d>0 are predefined parameters.

[0030] Regarding the first and second subsets of points in the frontal, left and right regions: these can be all the points Mj for which -Qo <Qj <0 (resp. O<07<0o )et-0o <<b / <+ o- Alternatively, it could be points with a direction circumscribed in a circular or elliptical cap shape.

[0031] As will be seen later, the distance of the ASk source can also condition the application of the flattening step. In other words, the front left and right regions can be defined independently of the distance of the source composing the audio scene to be reproduced, or alternatively, the front left and right regions can have a certain limited depth, as will be seen below, forming angular sectors finite in depth.

[0032] It is noted that the calculations can be done in linear or logarithmic mode using decibels in dB.

[0033] According to one embodiment, the flattening step (SAP) is performed in real time. An "on-line" switching is performed, that is to say, the filter is bypassed for sources that are located inside the front flattening regions.

[0034] According to one embodiment, the flattening step (SAP) is performed in offline mode (in practice, this is referred to as 'OFFLINE' mode). The filter values ​​in the memory storage area are replaced with those for the directions located within the frontal flattening regions. The modified filters are then used for the real-time binaural engine.

[0035] According to one embodiment, the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

[0036] For example, in practice, the flattening step is carried out for distances less than a predetermined threshold radius Rma, and the flattening step is not carried out for distances greater than this predetermined threshold Rma.

[0037] According to one embodiment, the process provides for smoothed joints at the edge of the front regions, respectively left and right.

[0038] The smoothing in question can be carried out in time (temporal smoothing) or in space (transition band).

[0039] In the configuration where flattening is performed only on the amplitude, this solution proves effective; in other words, a so-called "cross fade" process is used. Combing effects are avoided at the edges of the flattened front regions. The blending may involve, on a boundary strip, an interpolation between the filter values ​​just outside the flattened area and the flattening constant.

[0040] According to one embodiment, a further step, denoted SD, may be provided for evaluating the intraaural delay, denoted t(0, ¢), the average intraaural delay being calculated as follows:

[0041] =j^argmax(\^

[0042] J^argtnax[ | )

[0043] 7-(0)=^(^(0)- / ,(40))

[0044] fs being the sampling frequency of the HRIRs, and argmax(un) being the index of the maximum of a sequence un, * being the convolution product, w being a convolution filtering function, with time application on the impulse response or application in the frequency domain, the filter then being calculated as follows:

[0045] HHj ( ÿ = ] 0WW)) / 20

[0046] 0 / / 5(^04)) / 20^

[0047] Correction by intraural delay is particularly relevant in the case of flattening only on the amplitude.

[0048] According to one embodiment, it may further provide for a step of individualizing the intraural delay (SG) as a function of a head radius b of a target listener, the intraural delay being determined by with ^lc new head radius and ale original radius.

[0049] According to one embodiment, an angular correction step (SH) may also be provided, with a parameter y that can be adapted to each user, according to the following formulation: / (0) = (^+^(20),0)

[0050] According to one embodiment, a further step, denoted SB, may be provided, for smoothing the spectral amplitude by binaural relative deviation, and normalization for the reference direction (Oq, with mean over the subjects, according to the formulation: detaxed, ¢) ) =iE;v(r(M) - / 2(¾. ) dB(Hr( b, ¢) ) = e, ¢)¢) ) e, ¢) )

[0051] The invention also relates to an audio apparatus or device configured to carry out the steps of a process as described above. Depending on the configuration, the audio apparatus or device may (or may not) perform the flattening step.

[0052] A head tracking function may be provided that supplies, in real time, a signal denoted HDT(t) representing the head movement in yaw rotation, as well as in left-right lateral tilt and front-back tilt, thus three spatial rotation coordinates. The relative positions of the scene sources are corrected by compensating for the detected movements and displacements.

[0053] The invention also relates to a processing system configured to implement the steps of a process as described above.

[0054] The invention will be further detailed by describing non-limiting embodiments, and based on the accompanying figures illustrating variants of the invention, in which: - [Fig.1] is a schematic diagram illustrating the spherical coordinates of points and directions in space around an individual, for example a listener in the context of listening or in the context of measuring auditory transfer functions HRTF; - [Fig.2] is analogous to [Fig.1] in the case of an audio scene rendering and illustrates in particular an example of a right angular front region where the filters are flattened for the right ear filters; - [Fig.3] is analogous to [Fig.2] and illustrates in particular an example of left angular frontal region where the filters are flattened for left ear filters; - [Fig.4] is analogous to [Fig.3] and illustrates a particular example of left angular frontal region where the filters are flattened for left ear filters, the frontal region having a capped / limited depth; - [Fig.5] schematically illustrates an audio scene reproduction using a binaural rendering engine, for a single ASi audio source; - [Fig.6] illustrates an example of a frequency transfer function, for the right ear, with the amplitude and phase for four different azimuth positions and for a given elevation position; - [Fig.7] illustrates the example of the transfer function of [Fig.6] having undergone the flattening step only in amplitude; - [Fig.8] illustrates the example of the transfer function of [Fig.6] having undergone the amplitude and phase flattening step; - [Fig.9] schematically illustrates an example of the sequence of steps of the proposed process according to a first variant of hardware implementation; - [Fig. 10] schematically illustrates an example of the sequence of steps of the proposed process according to a second variant of hardware implementation; - [Fig. 11] schematically illustrates a more complete example of the sequence of steps of the proposed process according to a third variant.

[0055] In the various figures, the same reference numerals designate identical or similar elements. For the sake of clarity, some elements are not necessarily shown to scale.

[0056] Generalities and system

[0057] Generally, spatialized binaural sound reproduction allows a listener U to be immersed in an audio scene and to perceive sound sources coming from different directions and positions in space, whereas listener U only has two speakers HPg and HPd.

[0058] In the example of [Fig.5], the playback device is an audio headset 2 with the left speaker HPg arranged opposite the left ear OG of the listener and the right speaker HPd arranged opposite the right ear OD of the listener U.

[0059] In other configurations, the rendering organ can be a pair of intra-aural earbuds such as, for example, “earbuds”.

[0060] In yet another configuration, in a so-called open-air listening context sometimes referred to as 'transaural™', two reproduction devices, such as conventional loudspeakers, are arranged with crosstalk cancellation. The proposed invention can also be applied in this transaural configuration, with two or more loudspeakers.

[0061] For the different configurations mentioned above, the two speakers can be controlled by a single control unit directly in wired mode, or the speakers can be controlled by a single control unit but via a wireless Bluetooth connection.

[0062] In one scenario, the audio scene to be reproduced may consist of a single sound source, denoted AS i in [Fig. 5]. In another scenario, the audio scene to be reproduced consists of a plurality of sound sources AS j. For example, the audio scene to be reproduced may include NbS sound sources ASk (k ranging from 1 to N).

[0063] Each sound source to be reproduced, such as the first source AS i, is located at a position Ml characterized by its polar coordinates 01 and ¢1 and its distance RI (also denoted rl), as illustrated in [Fig. 1]. Polar coordinates will be used throughout this document. The first coordinate of a generic point M 0 represents the azimuth with respect to a "straight ahead" direction, with positive angles being oriented to the right. The second coordinate (which can also be noted) <j>) corresponds to the elevation, also called "site", relative to a horizontal direction, with positive angles being pointed upwards.

[0064] In [Fig. 1], the plane labeled PH is the horizontal reference plane for the listener with their head upright. The plane labeled PMS is the median sagittal plane, separating the left hemisphere from the right hemisphere. The plane labeled PVI is the vertical interaural plane (passing through 0° = 90°). When the system is equipped with the head-tracking function, the deviation of the actual position from the upright head position is compensated by the same angle values, but with opposite signs. For the listener, the sound source always appears to be in its given position in space, regardless of any small head movements the listener may make.

[0065] The reproduction of an audio scene is based on the use of acoustic transfer functions of the head between the positions of space (0, <e>) and the ear canal, called Auditory transfer functions. These transfer functions, called "HRTFs" (for "Head Related Transfer Functions"), concern the frequency form of the transfer functions. Their time-domain form, that is, the impulse response, will be referred to hereafter as "HRIR" (for "Head Related Impulse Response").

[0066] In the context of binaural playback, the following transfer functions are then used: a frequency transfer function for the left ear (HRTFg(M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRIRg(M)(t)) and an impulse response function for the right ear (HRIRd(M)(t)), M being the position of the source in polar coordinates, f the frequency and t the time.

[0067] Transfer functions are also referred to as 'filters' in this document.

[0068] The filters that will be used in the process described below may already be Whether symmetrical or not, they may have been averaged across N individuals, or they may originate from a single individual. The filters in question may be public or proprietary.

[0069] Frequency transfer functions are manipulated in this document in digitized / digital form, namely as a sequence of values. For a large number of frequencies within the human audible spectrum, the transfer function provides an amplitude and a phase. Typically, 128 frequency points may be listed, or even 256. For each point, the amplitude and phase are encoded in 8, 10, or 12-bit digital format.

[0070] The transfer functions can be coded in the 'sofa' format known in the art of sound engineers.

[0071] Returning to [Fig. 5], the audio scene is rendered by a binaural rendering engine. The binaural rendering engine has at its disposal the transfer functions processed in their frequency or time form. Thanks to these functions, for each sound source AS j of the audio scene to be reproduced, characterized by the wave time pattern ASi(t) of the sound source as well as its polar coordinates (0i, <e>i) and its distance from the listener Ri, the binaural rendering engine generates the corresponding three-dimensional audio scene in both speakers.

[0072] Fig. 5 represents only one sound source (01, C>l, Rl) but the reader will understand that what follows will be repeated for each of the sound sources, and that the respective signals of the different sound sources are added together in the waves produced towards the loudspeakers (HPg, HPd).

[0073] When the binaural rendering engine operates in the frequency domain, a Fast Fourier Transform (FFT) denoted XASi(f), more simply denoted XAS in [Fig. 5], is calculated from the time pattern ASi(t). The coordinates (01, l,rl) of the sound sources are applied on one side to the HH filter ] for the left ear and on the other side to the HH filter r for the right ear.

[0074] HH i and HH r are the filters resulting from upstream processing, with flattening where applicable, or flattening being directly performed by the binaural rendering engine.

[0075] For the left channel, the fast Fourier transform XAS of the source signal is multiplied by the filter HH ] assigned the coordinates of the source, which gives the signal to be reproduced in frequency form WASg. This signal is subjected to an inverse Fourier transform calculation 21, thus obtaining the time signal WASg(t), i.e., the background sound image, which will be played on the left loudspeaker HPg.

[0076] For the right channel, a similar procedure is used: the fast Fourier transform of the source signal is multiplied by the HH filter, assigned the coordinates of the source, which gives the signal to be reproduced in frequency form, WASd. This signal is then subjected to an inverse Fourier transform calculation, resulting in the time-domain signal WASd(t), i.e., the background sound image, which will be played on the right loudspeaker HPd.

[0077] In an alternative implementation, when the binaural rendering engine is working in the time domain, for the left channel a convolution product is performed between the sound source signal ASi(t) and the impulse response function hh / t) with coordinates (01, <bl,rl) de la source sonore on obtient ainsi le signal temporel WASg(t) qui va être joué sur le haut-parleur gauche HPg.

[0078] A similar procedure is used for the right channel; a convolution product is performed between the sound source signal ASi(t) and the impulse response function hh r(t) with coordinates (01, <bl,rl)de la source sonore on obtient ainsi le signal temporel WASd(t) qui va être joué sur le haut-parleur droite HPd.

[0079] It should be noted that to a certain extent the left and right transfer functions (HH ] and HH r) can be customized for the listener experiencing the three-dimensional audio scene. This relates to adapting the head diameter to internal delays, azimuthal angular correction ('gamma' correction), or combined azimuthal and elevation angular correction.

[0080] Finally, regarding [Fig.5], if the head tracking function is present, it provides in real time a signal denoted HDT(t) representing the movement of the head in yaw rotation, as well as in left-right lateral tilt and front-back tilt, therefore three spatial rotation coordinates. As explained elsewhere, the rendering engine compensates by injecting these values ​​with an opposite sign into the sound signal generation algorithm.

[0081] Advantageously, according to the present invention, the filters are flattened or flattened. This is illustrated in particular in [Fig.2] and [Fig.3].

[0082] A right frontal region ZFd is defined, visible in [Fig.2], in which the amplitude and / or phase of the right ear filters are set to a constant value.

[0083] Furthermore, a left frontal region ZFg is defined, visible in [Fig.3], in which the amplitude and / or phase of the left ear filters are set to a constant value.

[0084] The left frontal region ZFg encompasses a first subset of points Mj for which -00 < 0j < 0 and - <F0< <bj <+<h0. La région frontale droite ZFd englobe un deuxième sous ensemble de points Mj pour lesquels O<0j<0o et On recall here that azimuth angles 0 are expressed with positive values ​​on the right side, and negative values ​​on the left side, 0 = 0 being the direction straight ahead.

[0085] However, the inverse reference for the sign of 0 could be used, mutatis mutandis.

[0086] The step of defining the frontal regions ZFg and ZFd is marked DRF in Figures 9 to 11

[0087] The flattening performed may only concern the flattening of the filter's amplitude without affecting the phase portion. In an alternative embodiment, the flattening step concerns both the amplitude and the phase of the filter.

[0088] Regarding the first and second subsets of points in the frontal, left and right regions, these can be all the points Mj for which -00 < 0j < 0 (resp. 0 < 0j < 00) and - <ho<<I)j<+ o, in which case we have a rectangular region as illustrated in the figures.

[0089] Alternatively, the first and second subsets of points may concern points with a direction circumscribed in a circular or elliptical cap shape, as illustrated by the dotted line ZF'.

[0090] The constant value used in the flattening step applies to all frequencies of the audio spectrum (either for amplitude only or for amplitude and phase as discussed previously).

[0091] In a particular case, this value is zero, in other words, a complete selective bypass of the filter is carried out in the front, left and right regions.

[0092] The reference angles 0O and <h0 sont paramétrables. Les angles de référence 0O , <e>These can correspond to 00 = 15° and O0 = 15°. These are referred to as 'image' filters in the context of the present invention. According to another solution, the reference angles Qo^ocan correspond to 0O =30° and <F0=15o. Il s’agit de filtres dit ‘audio’ dans le contexte de la présente invention.

[0093] According to a basic embodiment, the flattening of the filter is applied to the aforementioned front regions without taking into account the distance from the audio source to be reproduced.

[0094] However, according to one option of the method, the distance of each audio source can play a role in the filter flattening process. In other words, the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

[0095] In practice, turning to [Fig.4] which illustrates this configuration, the flattening step is carried out for distances less than a predetermined threshold radius Rma, and the flattening step is not carried out for distances greater than this predetermined threshold Rma.

[0096] As an alternative to a binary logic with respect to the distance from the source, it is not excluded to use a progressive modulation, e.g. we flatten less and less as the distance increases.

[0097] In Figures 6 to 8, which illustrate filters, the x-axis represents frequency, graduated logarithmically. The y-axis represents amplitude and phase in decibels, respectively.

[0098] Figure 6 illustrates an example of a frequency transfer function, for the right ear, with the amplitude and phase for four given azimuth positions and for one azimuth position, without a flattening step.

[0099] Curve 81 (solid line) represents the amplitude of the frequency transfer function, for the right ear, for 0 = 0°.

[0100] Curve 82 (small dotted line) represents the amplitude of the frequency transfer function for 0 = 15°. Curve 83 (dashed line) represents the amplitude of the frequency transfer function for 0 = 30°. Curve 84 (solid line) represents the amplitude of the frequency transfer function for 0 = 45°.

[0101] Curve 85 represents the phase of the frequency transfer function for the right ear at 0°. Curve 86 (small dotted line) represents the phase of the frequency transfer function at 0°. Curve 87 (dashed line) represents the phase of the frequency transfer function at 0°. Curve 88 represents the phase of the frequency transfer function at 0°.

[0102] Here, ¢=0° has been taken as the reference for the filter examples.

[0103] Figure 7 reproduces the example of the transfer function from Figure 6, with the flattening step applied only in amplitude. According to the illustrated example, the reference angle is 00 = 35°.

[0104] For the amplitude, curves 81 to 83 are replaced by a flat curve 91, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. On the other hand, curve 84 for 0 = 45° is not modified and remains unchanged (direction outside ZFd).

[0105] The phase curves remain unchanged.

[0106] [Fig.8] reproduces the example of the transfer function of [Fig.6] with application of the flattening step in amplitude and in phase.

[0107] For the amplitude, curves 81 to 83 are replaced by a flat curve 91, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. However, curve 84 for 0 = 45° is not modified and remains unchanged.

[0108] For the phase, curves 85 to 87 are replaced by a flat curve 92, as they concern directions / orientations that are encompassed within the frontal flattening region ZFd. On the other hand, curve 88 for 0 = 45° is not modified and remains unchanged.

[0109] If for the same filters we took as reference angle 0O = 22°, the curves 81 to 82 would be replaced by the flat curve 91, but the curves 83 and 84 would remain unchanged (directions outside ZFd).

[0110] The method advantageously provides for smoothed connections at the location of the border of the front regions ZFg and ZFd, respectively left and right.

[0111] The smoothing in question can be carried out in time (temporal smoothing) or in space (transition band).

[0112] The different stages of reprocessing these transfer functions are illustrated in [Fig.9] to IL The first stage noted S0 corresponds to the provision of the above-mentioned transfer functions from databases or measurements with, where appropriate, right-left symmetrization, and where appropriate, averaging over N individuals.

[0113] In the generic example of Figures 9 and 10, the step of defining the DRF flattening front areas is followed by the step of flattening the SAP filters. In this example, there is no other correction, and the result of the flattening is used to run the binaural rendering engine (SR step).

[0114] According to the first possibility, illustrated in [Fig. 10], the SAP flattening step is carried out in real time. An "on-line" switching is performed, i.e., the filter is bypassed for sources located within the front flattening regions.

[0115] According to the second possibility, illustrated in [Fig. 9], the SAP flattening step is performed in offline mode. The filter values ​​for the directions located at are then replaced in a memory storage area. the interior of the frontal flattening regions. The filters thus modified are then used for the real-time binaural engine.

[0116] Furthermore, a spectral amplitude symmetrization step SA may be provided, including additional smoothing by local average of k neighboring directions j, according to the formulation: e, ¢) = wMQj, ¢.) ) +dB(wJHii-) J \ \ \ J ' / x 3 J f / / 101171 ff (M) = î ( <» ( >^ fjHÿ - e f ¢.) ) )

[0118] According to one option, the method provides for smoothing the spectral amplitude by binaural relative deviation, normalization for a reference direction (¢Q) with average over the subjects, according to the formulation, for a source on the left: 101191 101201 dB(H,(e, » ) = ^(11(8, ¢)-L'(8,i) ) + dB(Hl(8, ¢) )

[0121] This step is denoted SB. For a right-hand source, normalization with respect to the 00.3¾ right-hand reference direction can be applied similarly: 101221 dB(HM ¢) )=41^(^(8, ¢))

[0123] dB(H[( 0,( / ))) = (r< M) -#(M)) +dB(Hr( 0,( / )))

[0124]

[0125] It is noted that the flattening of the filters on the front regions ZFg ZFd can also be expressed mathematically according to the formulation written differently:

[0126] dB^B, ¢) ) = ^(^(0, ¢)-Ü( 10\, ¢))

[0127] dB(Hr(0, ¢) ) U ¢) -L\0, ¢) ) + dB(Hl( 0, ¢) )

[0128] Furthermore, a step denoted SC of symmetrization and phase averaging can be carried out, according to the formulation: * \ •. d J / / 101291<Hr(M)> = 41 / (<H‘J(e,4)> + > )

[0130] The notation < . > denotes the phase of the complex transfer function.

[0131] Furthermore, a step denoted SD can be carried out, evaluating the mean intra-aural delay denoted t(0, ¢), the mean intra-aural delay r(6, ¢) being calculated as follows; / / (0, ¢) = argmaxi 0, $) )

[0132] ) - -1 ^argmax( |Hd( 0, ÿ | ) 101331 ^¢)=4(1,(8,^)-1,((^))

[0134] fs being the sampling frequency of the HRIR impulse responses, and argmax(Un) being the index of the maximum of a sequence un, of digitized values ​​of the impulse response, the notation * being the convolution product, w being a convolution filtering function. The convolution filtering function w can be reduced to a Dirac pulse (then we simply select the largest signal peak).

[0135] The mean intraural delay can be applied in the time domain, by applying a corresponding delay to the impulse response.

[0136] The average interaural delay can be applied in the frequency domain, the filter then being calculated as follows: [one / ] =

[0138] HHr( $ y „

[0139] Furthermore, we can proceed to an individualization step of the interaural delay denoted SG, as a function of a head radius b of a target listener, the interaural delay being determined by |T, with b the new head radius and a the original radius.

[0140] The radius may have been derived from measurements accompanying the transfer function measurements. The radius may have been derived from anthropological data.

[0141] The radius b can be chosen according to the listener benefiting from the audio scene rendering. For example, the radius b can be smaller than a for a listener with a small head circumference; we would refer to size XS or S.

[0142] For example, radius b may be larger than radius a for a listener with a large head circumference, we will speak of size L or XL.

[0143] For example, radius b may be larger than radius a for a listener with an average head circumference.

[0144] Then we proceed to a step denoted SH of angular correction.

[0145] According to a first simple example referred to here as 'in two dimensions', a correction is made azimuthal angular, with a parameter y that can be adapted to each user, according to the following formulation:

[0146] f(6, ¢) = (0+y*sin(20), ¢)

[0147] Put another way, instead of using the azimuth angle 0 directly, we use instead a transform of the azimuth angle in the form 0 + y^sin ( 20 ).

[0148] The value of y is configurable.

[0149] If the 'head movement tracking' function is implemented, the SH step performs yaw movement corrections on 0 in a manner similar to that shown above.

[0150] According to another example of the SH angular correction step, correction is made in azimuth and elevation, according to a so-called 'three-dimensional' deformation, using f ( ¢) = ( 0+y*sin(20), 0 + y'*sin(0)) • The parameters y and y' can be adapted to each user.

[0151] In an even more generic way, the step denoted SH of angular correction consists of replacing the angles 0,q> in the HRTF formulas by the angles (0^) such that (O',^) = f (0,0) f being any azimuth and elevation correction function.

[0152] Figure 11 shows a schematic of the process with the aforementioned options. The input metadata, namely the initial filters / transfer functions, and the definition of the right and left front zones (steps S0 and DRF), are shown at the beginning. The subsequent steps, namely SB, SAP, SD, SF, and SG, are not necessarily applied or performed in the order presented.

[0153] Other considerations

[0154] Optionally, due to the symmetrization, it is possible to keep the filters in memory only for one hemisphere, right or left, and to use a planar symmetry with respect to the PMS sagittal plane performed by the real-time binaural rendering engine.

[0155] More precisely, in practice, after the phase and amplitude symmetrization steps SA and SC, data can be stored for only one hemisphere, and the complementary transformations performed only on that hemisphere. The process for the other hemisphere will be reconstructed in real time by planar symmetry with respect to the sagittal plane PMS. In this way, the memory space occupied can be halved, which may be relevant for certain categories of audio hardware or processing chains.

[0156] According to another optional feature, in addition to the transformations described above, a further transformation may be provided consisting of collectively boosting high-pitched sounds across all filters. Thus, for example, sounds with frequencies above 2 or 3 kilohertz may be boosted, with this boost having an amplitude between 1 dB and 3 dB, for example.< / e> < / e> < / e> < / j> < / e> < / e> < / e> < / e> < / e> < / e> < / e>

Claims

1. Demands A method for generating an audio scene in a binaural spatialization system, the method comprising: S0- a step of making available a set of primary auditory transfer functions, relating to the left and right ear of individual(s), the set of primary auditory transfer functions comprising for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(Mj)(f)) denoted Hgtdpdr) and comprising amplitude and phase, a frequency transfer function for the right ear (HRTFd(Mj)(f)) denoted Hd(0j, <bj) et comprenant amplitude et phase, une fonction de réponse temporelle impulsionnelle pour l’oreille gauche (HRIRg(Mj)(t)) notée hg(0j,<bj), et une fonction de réponse temporelle impulsionnelle pour l’oreille droite (HRIRd(Mj)(t)) notée hd(0j,<bj), chaque position Mj ayant pour direction en coordonnée polaires azimut Qj et élévation <b,, les fonctions de transfert fréquentielle et temporelle étant appelées généralement filtres,which may also be the result of an averaging operation based on results from several individuals, and which may or may not be symmetrical, the audio scene to be generated comprising one or more audio sources ASk, each located at a distance Rk, an azimuth 0k and an elevation <Fk, le procédé étant caractérisé en ce qu’il comprend :, - the definition of a left frontal region (ZFg) comprising a first subset of points Mj for which -00 <07 <0 and - <f>0 < ¢ / <+ <l\ - the definition of a right frontal region (ZFd) comprising a second subset of points Mj for which O<07<00 and - 0 < <h <+<ty- a filter flattening step (SAP) in which the amplitude and / or phase of the left-ear filters are set to a constant value over the left frontal region, and the amplitude and / or phase of the right-ear filters are set to a constant value over the right frontal region, - a usage step (SR) to generate an audio scene rendering using a binaural rendering engine based on filters with flattening and the audio sources of the audio scene to be generated.

2. A method according to claim 1, wherein the flattening step (SAP) is carried out in real time.

3. A method according to claim 1, wherein the flattening step (SAP) is carried out in offline mode.

4. A method according to any one of claims 1 to 3, wherein the application of the flattening step is conditioned and / or modulated by the distance Rk of each source ASk composing the audio scene to be reproduced.

5. A method according to any one of claims 1 to 4, wherein smoothed joints are provided at the edge of the front regions, respectively left and right.

6. A method according to any one of claims 1 to 5, wherein a step denoted SD is provided for evaluating the average intraaural delay denoted t(0, ¢), the average intraaural delay being calculated as follows: / z(ft ¢) = 77^ argmax{ \h'g(B, ¢) ) ÿ) = argmaxi ^)**l ) 7-(0,0)=^(0.0)- / / (0.0)) where fs is the sampling frequency of the HRIRs, etargmax(un) is the index of the maximum of a sequence un, * is the convolution product, w is a convolution filtering function with a time-domain application on the impulse response or an application in the frequency domain, the filter then being calculated as follows: HH^B, ¢) = HHr(^ ¢) = 1

7. A method according to any one of claims 1 to 6, further comprising a step of individualizing the intraural delay (SG) as a function of a head radius b of a target listener, the intraural delay being determined by with b the new head radius and a the original radius.

8. A method according to any one of claims 1 to 7, further comprising an angular correction step, denoted SH, with a

9.

10. parameter y can be adapted to each user, according to the following formulation: / (0,^) = (0 + y*sin(20), ¢) Audio device configured to implement the steps of a process according to any one of claims 1 to 8. Processing system configured to implement the steps of a process according to any one of claims 1 to 8.< / h> < / f>