Method for generating an audio scene in a binaural spatialisation system
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2026-04-08
AI Technical Summary
Existing binaural spatialization systems fail to accurately generate three-dimensional audio scenes for a general audience, as using individual-specific auditory transfer functions or averaged functions do not fully satisfy subjective results, and existing methods do not adequately restore timbre or eliminate comb filter effects.
A method that involves providing a set of primary auditory transfer functions for each individual, including frequency and impulse response functions for both ears, with steps of symmetrization, normalization, and reprocessing to create reworked transfer functions that are then used in a binaural rendering engine to generate an audio scene, which includes steps like transition to logarithmic values, smoothing, and convolution operations.
The method effectively reproduces the timbre of audio sources for a diverse group of users, eliminates comb filter effects, and provides improved balancing and perceived quality, making it suitable for a wide range of audio scenarios, including virtual reality and video games.
Smart Images

Figure EP2024064427_28112024_PF_FP_ABST
Abstract
Description
[0001] Method for generating a scene in a binaural spatialization system
[0001] The invention relates to a method for generating an audio scene in a binaural spatialization system. The invention also relates to a system and an audio headset in which the proposed method can be implemented.
[0002] Generally speaking, it is increasingly common to need to synthesize a three-dimensional audio scene, whether for listening to stereo music on headphones or via earbuds or earplugs, or in the case of immersion scenarios in virtual reality with a virtual reality headset.
[0003] Binaural hearing occurs when both ears of an individual receive specific sound signals, corresponding to a perception of 15 sounds in space allowing the auditory system and the individual's brain to determine the direction of origin of a sound.
[0004] Thus, thanks to binaural hearing, spatialized sound reproduction allows a listener to perceive sound sources coming from any direction or position in space, and to have the impression of being immersed in a three-dimensional sound scene.
[0005] The signals received at the listener's right and left ears are differentiated, notably but not exclusively, by their intensity and by the arrival and phase delay.
[0006] The term "multichannel", in processing for spatialized sound reproduction, consists of producing a representation of an acoustic scene in the form of P signals (called spatial components). These signals contain all the sounds that make up the sound scene, but with weightings that depend on their direction (or "incidence") and described by P associated spatial rendering functions. It should be noted, however, that the proposed method also works for a scene with a single audio source.
[0007] The particular techniques of spatialized sound reproduction to which the present invention relates are based on the existence of acoustic transfer functions of the head between the positions in space and the auditory canal. These transfer functions, called "HRTF" (for "Head Related Transfer Functions"), concern the frequency form of the transfer functions. Their temporal form will hereinafter be designated "HRIR" (for "Head Related Impulse Response").
[0008] More particularly, we are particularly interested in a set of auditory transfer functions comprising for a given individual and for a plurality of measured positions M: a frequency transfer function for the left ear (HRTFg(M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRIRg(M)(t)) and an impulse response function for the right ear (HRIRd(M)(t)).
[0009] There are public databases in which these auditory transfer function measurements have been listed for a plurality of individuals. In the remainder of this document, a particular individual will be assigned the index 'i' and the number of individuals will conventionally be equal to N.
[0010] The so-called ipsilateral ear corresponds to the ear that receives the sound first, and the so-called contralateral (or 'contralateral') ear corresponds to the ear that receives the sound second for the same sound source. In the rest of the document, 10 we call the left hemisphere the half-sphere located to the left of the sagittal plane referenced by the head of the listener in an upright position, and the right hemisphere the half-sphere located to the right of this sagittal plane.
[0011] The left ear is the ipsilateral ear for sounds originating from the left hemisphere and the right ear is the ipsilateral ear for sounds 15 originating from the right hemisphere.
[0012] To generate a sound scene by synthesis in a binaural rendering engine, a convolution operation is carried out between the sound sources of the scene and the temporal impulse responses, respectively right and left, or if working in the frequency domain, 20 a multiplication operation of the spectra then an inverse Fourier transform.
[0013] It turns out that using a set of auditory transfer functions measured on a particular individual is not satisfactory if used for a general binaural rendering engine used for a general audience. 25
[0014] Building an average of the measurements on N individuals and producing a set of average auditory transfer functions is also not entirely satisfactory.
[0015] Some have tried to work on the phase of transfer functions as for example in document US10609504 but the subjective results 30 still leave something to be desired.
[0016] The inventors sought to improve the situation, notably by reworking the transfer function sets from public database(s), with a particular target of binaural rendering in headphones. Particular attention is paid to the reproduction of the timbre of audio sources, 35 and to the elimination of possible comb filter effects.
[0017] The proposed method and system can also be used in a so-called “transaural” context with listening on loudspeakers.
[0018] The inventors also kept the objective of being able to offer a possibility of individualization with regard to the listener who will be immersed in the three-dimensional audio scene.
[0019] To this end, a method is therefore proposed here for generating an audio scene in a binaural spatialization system, the method being characterized in that it comprises: S0- a step of providing a set of primary (basic) auditory transfer functions, relating to the left and right ears of individuals, taken from a set of N individuals, the set of primary auditory transfer functions comprising for each individual i and for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(i,Mj)(f)) noted ^^ ^ ൫ ^^ ^ , ^^ ^ ൯ and including amplitude and phase, a frequency transfer for the right ear (HRTFd( ௗ ^ i,Mj)(f)) noted ^^ ൫ ^^ amplitude and phase, a temporal impulse response function for the left ear (i,Mj)(t)) noted ℎ ^ ^ ൫ ^^ ^ , ^^^ ൯ and a time impulse response function for the right ear (HRIRd(i,Mj)(t)) denoted ℎ ௗ ൫ ^^ ^ , ^^ ^ ൯, each position Mj having as direction in polar coordinates ^j and Φj (azimuth and elevation), the method comprising: SA- a preliminary preparation step noted SA, in which we a passage into logarithmic values, with where appropriate (if this has not already been done in the public or private primary base) a symmetrization of the spectral amplitude for all directions ( ^^, ^^) of a subject ^^, using an averaging function FA, according to the formulation: ^^ ^ ( ^^, ^^) = ^^ ^^ ^ ^^ ^^ ( ^^ ^ ^ ^ ^ ( ^^ ^^ , ^^ ^^ )) , ^^ ^^ ( ^^ ^ ^ ^ ^ (− ^^ ^^ , ^^ ^^ ))^ ^ the dB FA function can be in a case SB- a step noted SB, of smoothing the spectral amplitude by binaural relative deviation, normalization for a reference direction ( ^^ ^ , ^^ ^ ) with average on the subjects, according to the formulation: ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = 1 ே ^ ^ ^ ^ ^^ ^ ( ^^) − ^^ ^ ൫ ^^0, ^^ ൯^ D- a ; ^^ ^ ^^ = ^^ ^ ^^ ^^ ^^ ^^ ^^ ^^൫หℎ ^ ^ ^^ ∗ ^^ห൯ being the HRIR, and ^^ ^^ ^^ ^^ ^^( ^^ ) being values e the convolution response, w being a convolution filter function, which can be reduced to a Dirac impulse when we simply select the largest signal peak), SF- a step denoted SF, of a set of reprocessed transfer functions, denoted HHl and HHr, ^^ ) SR- a step s ^^ ^^ ^ ^^, ^^ ^^ ^^ ( ^^, ^^) for to a m ( engine ) binaural rendering, with one or more point audio sources (s) prescribed.
[0020] Thanks to the provisions presented above, the development and use of the reprocessed transfer functions gives very satisfactory results. The timbre of the audio source or audio sources is restored in a very satisfactory manner for a set of users in a tested panel.
[0021] It is further noted that the normalization of the amplitudes of spectral filters, with respect to a reference direction ൫ ^^0, ^^ 0൯ proves to be particularly beneficial, especially from the point of view of general balancing.
[0022] Advantageously, we note that the comb filter effects are erased.
[0023] The average effect in relation to all the individuals measured makes it possible to work on an average individual and to maximize the adequacy of the reworked transfer functions in relation to the general population able to use devices implementing the proposed process.
[0024] It is noted that the number of audio sources participating in the generation of synthesized audio signals is arbitrary. A single audio source may correspond to a case of a video conference with a person delivering a speech in a location without any other sound source in the audio scene. In other configurations, a multiple sound environment is provided with multiple sound sources located in front of and behind the listeners, as in the case of immersion in a subjective view video game (for example, FPS type for 'First Person Shoot') or as in the case of simulators of all types.
[0025] It should also be noted that according to the proposed method, we switch to logarithmic values to perform calculations and then conversely we switch back to linear values before injecting the filters with their amplitudes and phases into the binaural rendering engine. This allows us to do subtractions instead of divisions and this lightens the overall computational load.
[0026] As will be specified later, certain steps are carried out in a so-called offline mode, i.e., before actual use in a binaural rendering engine in an audio rendering device, whereas conversely, certain steps are carried out in real time, i.e., at the time of the generation of sound waves in an audio rendering device. As we will see, there are several possible solutions for the functional division between the offline steps and the real-time steps.
[0027] Even more generally, a method is therefore proposed here for generating an audio scene in a binaural spatialization system, the method comprising: S0- a step of providing a set of primary (basic) auditory transfer functions, relating to the left ear and the right ear of individuals, taken from a set of N individuals, the set of primary auditory transfer functions comprising for each individual i and for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(i,Mj)(f denoted ^ )) ^^ ൫ ^^ , ^^ ൯ and including amplitude and phase, a frequency transfer for the right ear (HRTFd(i,Mj)(f)) noted ^ ^ e ^ ൫ ^^ , taking amplitude and phase, a time impulse response function for the left ear (HRIRg(i,Mj)(t)) noted ℎ ^ ^ ൫ ^^ ^ , ^^ ^൯ and a time impulse response function for the ore (HRIRd(i,Mj)(t)) denoted ௗ ^ right city ℎ ൫ ^^ ^ , ^^ ^ ൯, each position Mj direction in polar coordinates ^j and Φj (azimuth and elevation), the : SA- a preliminary preparation step noted SA, in which we a passage into logarithmic values, with where appropriate (if this has not already been done in the primary, public or private base) a symmetrization of the spectral amplitude for all directions ( ^^, ^^) of a subject ^^, using an averaging function FA, according to the formulation: ^^ ^ ( ^^, ^^) = ^^ ^^ ^ ^^ ^^ ( ^^ ^ ^ ^ ^ ( ^^ ^^ , ^^ ^^ )) , ^^ ^^ ( ^^ ^ ^ ^ ^ (− ^^ ^^ , ^^ ^^ ))^ ^ the dB FA function can be in a case SB- a step noted SB, of smoothing the spectral amplitude by binaural relative deviation, normalization for a reference direction ( ^^ ^ , ^^ ^ ) with average on the subjects, according to the formulation: ே ^^ ^^൫ ^^ ^ ^^, ൯ = 1 ^ ^ ^ ^ ^^ ^ ^^ − ^^ ^ ൫ ^^0, ^^ ൯^ SD- a noted ^^( ^^, ^^); SF- one of retired transfers, noted respectively ^^ ^^ ^ ( ^^, ^^) ^^ ^^ ^^ ^^ ^ ( ^^, ^^) ; SR- a step noted SR of using the reprocessed functions ^^ ^^ ^ ( ^^, ^^) and ^^ ^^ ^ ( ^^, ^^ to generate an audio scene in real time using a mote n ) bi-aural rendering core, with one or more point audio sources (s) prescribed.
[0028] As already mentioned above, normalizing the amplitudes of spectral filters, with respect to a reference direction ൫ ^^0, ^^0൯ turns out to be particularly beneficial, especially from a general point of view.
[0029] According to one aspect, at the SD- step, the average intraural delay denoted ^^( ^^, ^^) is as follows: ^ = 1 ே ^^ ^^ ^ ^^ ^^ ^^ ^^ ^^ ^^൫หℎ ^ ^ ∗ ^^ห൯
[0030] fs being the and ^^ ^^ ^^ ^^ ^^ ^^( ^^ ^ ) being of the , convolution, w being a filtering function of
[0031] We thus specify an example of a concrete way of calculating the average interaural delay.
[0032] According to one aspect, in step SF-, the set of reprocessed transfer functions is calculated as follows: ^^ ^^ ^ ( ^^, ^^) = 10ௗ^൫ு^(ఏ,థ)൯ / ଶ^^^
[0033] We have just calculated the set of functions of
[0034] According to one aspect, the method may further comprise a step denoted SB2 of flattening the amplitude of the filters on a frontal cap (CFg) for the ipsilateral ear, according to the formulation: ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = ^^ ^^ ^ ^^ ^ ൫ ^^0, ^^0൯^ for any direction ^j , Φj included in the respective angle [0, ^ 0, ] and [- Φ 0, +Φ0] left side. The source located on the left. If the source is located on the right then ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = ^^ ^^ ^ ^^ ^ ൫ ^^0, ^^0൯^ for any direction ^j , Φj included in the respective [0, ^0,] and [- Φ0, +Φ0] side
[0035] Advantageously, the amplitudes applied for flattening correspond to the amplitudes of the filter of the reference direction ^0, Φ0 of the ipsilateral side. According to one option, the filters for the contralateral ear remain unchanged, in particular concerning the amplitudes and the minimum phases.
[0036] Whereby, a flat filter without correction is applied to the amplitude of the filter of the ipsilateral ear, in a frontal cap. The test campaigns carried out by the inventors revealed that this arrangement was favorable from the point of view of perceived quality on the part of the users of the test panel.
[0037] We note that at the boundaries of the cap there is no jump or discontinuity, because the amplitudes have previously been normalized with respect to the reference amplitude in the direction [ ^0, Φ0]. The filter is continuous and differentiable over the entire range of angles.
[0038] According to another of this step, the step noted SB2 of the amplitude of the filters on a frontal cap (CF), can be according to the formula
[0039] ^^ ^^൫ ^^ ^ (^^, ^^)൯ = ^tion: ே ∑ ே ^ ^ ^^ ^ ( ^^, ^^) − ^^ ^ (| ^^|, ^^)^ ே 5 ൯ on the side of the to be applied symmetrically for a source located on the right.
[0041] According to one aspect, the binaural rendering engine further comprises a head tracking function. In practice, a 3-axis gyroscopic sensor provides the yaw, roll and pitch movements to an audio headset on which the binaural rendering engine is implemented. It should be noted here that the gyroscopic sensor can be integrated into the audio headset. Alternatively, the yaw movement can be determined by the use of one or more cameras, whether or not they are embedded in the headset. Thus, the rendering engine can compensate in real time for the listener's head movements by calling the corresponding inverse positions in the database: e.g., if a listening subject turns his head 15° to his left, a sound initially placed at 0° in front of him will shift to his right by 15°.
[0042] According to one aspect, the method may further comprise, before the SF step or during the SR step, a step denoted SE of interpolation of the spectral amplitudes, phases and interaural delays.
[0043] Whereby, starting from a base of known transfer function sets 25 with a fairly large step, for example every 15° or every 10°, it is possible to provide, thanks to the interpolation, a base of transfer function sets of better granularity, for example every 3° or even every 2°.
[0044] It should be noted that this interpolation step may be carried out in the offline steps or may be carried out in real time with the binaural rendering engine 30.
[0045] According to one aspect, the step of symmetrization of the spectral amplitude (SA) can further comprise smoothing by local average of k neighboring directions ^^, according to the formulation ^lation: ^^. ^ ^^ ^^ = 1 ^ ^^ ^^ ^ ^^ ^ ^^ ^ ^ ൫ ^^ ^ ^^ ^൯^+ ^^ ^^ ൫ ^^ ^ ^ ^ ൯^ 35
[0046] We also in space.
[0047] According to one aspect, the method may further comprise a step of individualizing the interaural delay (SG) as a function of a head radius b of a target listener, the interaural delay determined by ^^^^ ^^, with b the new head radius and a the original radius, i.e. the reference head radius.
[0048] According to one aspect, the method may further comprise a step denoted SH of angular correction, in fact an azimuthal angular correction, with a parameter γ that can be adapted to each user, according to the formulation: ^^( ^^, ^^) = ( ^^ + ^^ ∗ sin(2 ^^), ^^)
[0049] The parameter γ is preferably configurable, and may even be a function of the listener benefiting from the binaural audio rendering. In other words, instead of using the azimuth angle ^^ directly, a transform of the azimuth angle in the form ^^ + ^^ ∗ sin(2 ^^) is used instead. It turns out that the qualitative result is further improved by this angular deformation.
[0051] According to a related aspect, the method may comprise a step denoted SH of angular correction with an angular correction in azimuth and elevation 15 where the angles θ,φ in the HRTF formulas are replaced by the angles (θ′,φ′) such that (θ′,φ′) = f (θ,φ).
[0052] According to an example, one can choose (θ′, ^^′) = f( ^^, ^^) = ( ^^ + ^^ ∗ sin(2 ^^), ^^ + ^^. ᇱ ∗ sin( ^^)).
[0053] The parameters γ and γ' can be adapted to each user. 20
[0054] In other words, instead of using the azimuth angle ^^ directly, a transform of the azimuth angle in the form ^^ + ^^ ∗ sin(2 ^^) is used instead and instead of using the elevation angle φ directly, a transform of the elevation angle in the form ^^ + ^^′ ∗ sin( ^^) is used instead.
[0055] According to one aspect, the method may further comprise a step denoted 25 of symmetrization and averaging of the phase, according to the formulation: < ^^ ^ ( ^^, ^^) > = 1 ே 2 ^^ ^ ^ < ^^^ ^ ( ^^, ^^) > + < ^^ ௗ ^ ^(− ^^ ^^ , ^^ ^^ ^ > ^ ^
[0056] The complex. We which have 30 to the transfer function measurement campaigns. The invention also relates to an audio headset configured to implement steps SE to SR of a method as described previously. Depending on the configurations, the audio headset can carry out steps SG to SR or SH to SR. 35 The invention also relates to a processing system configured to implement steps S0 to SG of a method as described previously.
[0060] According to another embodiment of the present invention, starting from HRTF filter bases, possibly already symmetrized, all the flattening defined in step SB2 is applied. In this case, it may be necessary to connect the filters by continuity to the edge of the cap CFg, (respectively CFd) to guarantee jumping. This can be done by polynomial interpolation.
[0061] The invention will be further detailed by the description of non-limiting embodiments, and on the basis of the appended figures illustrating 5 variants of the invention, in which: - [Fig.1] is a schematic diagram illustrating the spherical coordinates of the points and directions of space around an individual, for example a listener in the context of listening or in the context of a measurement of the HRTF auditory transfer functions; 10 - [Fig.2] illustrates an example of frequency transfer function amplitude, with in addition the phase, for the contralateral ear with normalization with respect to the 15° direction; - [Fig.3] illustrates an example of frequency transfer function amplitude, with in addition the phase for the ipsilateral ear with 15 normalization with respect to the 15° direction; - [Fig.4] illustrates an example of frequency transfer function amplitude, with the phase for the contralateral ear with normalization with respect to the 30° direction, and flattening on the frontal cap; 20 - [Fig.5] illustrates an example of frequency transfer function amplitude, with the phase, with the phase for the ipsilateral ear with normalization with respect to the 30° direction, and flattening on the frontal cap; - [Fig.6] is analogous to Figure 1 and illustrates in particular areas of the frontal cap where the filtering spectra are flattened; - [Fig.7] schematically illustrates an audio scene rendering using a binaural rendering engine, for a single audio source AS. 1; - [Fig.8] schematically illustrates an example of a succession of steps of the proposed method according to a first hardware implementation variant; - [Fig.9] schematically illustrates an example of a succession of steps of the proposed method according to a second hardware implementation variant; - [Fig.10] schematically illustrates an example of a succession of steps of the proposed method according to a third hardware implementation variant.
[0062] In the different figures, the same references designate elements or similar. For reasons of clarity of the presentation, some are not necessarily represented to scale. 40
[0063] General information and system
[0064] Generally, a spatialized binaural sound reproduction allows one to perceive sources in an audio scene and directions and positions in space, while the listener U only has two loudspeakers HPg and HPd.
[0065] In the example of figure 7, the restitution is an audio headset with the left loudspeaker HPg arranged opposite the left ear OG of the listener and the right loudspeaker HPd arranged opposite the right ear OD of the listener U. 5
[0066] In other configurations, the restitution organ can be an in-ear pair such as for example "earplugs". In yet another configuration, in a so-called open-air listening context sometimes called 'transaural™', two restitution organs are provided, such as conventional loudspeakers, with cross-talk cancellation 10. The proposed invention can also be applied in this transaural configuration, with two or more loudspeakers.
[0068] For the different configurations mentioned above, the two loudspeakers can be controlled by a single control unit directly in wired mode, or the loudspeakers can be controlled by a single control unit but via a Bluetooth type wireless link.
[0069] In one case, the audio scene to be reproduced may consist of a single sound source, denoted AS1 in Figure 7. In another case, the audio scene to be reproduced consists of a plurality of sound sources ASi. For example, the audio scene to be reproduced may comprise NbS sound sources.
[0070] Each sound source to be reproduced, such as for example the first source AS1, is located at a position M1 characterized by its polar coordinates ^1 and Φ1 and its distance r1, as illustrated in Figure 1. The polar coordinates will be used throughout this document.The first coordinate ^ 25 represents the azimuth relative to a "straight ahead" direction, with positive angles being marked to the right or left depending on the side of the source concerned. The second coordinate Φ (also noted φ) corresponds to the elevation, also called "site", relative to a horizontal direction, with positive angles being marked upwards. 30
[0071] In Figure 1, the plane marked PH is the horizontal reference plane for the listener with a straight head. The plane marked PMS is the midsagittal plane; it separates the left hemisphere from the right hemisphere. The plane marked PVI is the interaural vertical plane (it passes through ^=90°). When the system is equipped with the head movement tracking function, the deviation of the actual position relative to the straight head position is compensated for by the same angle values by assigning them an opposite sign.For the listener, the sound source always appears to be at its given position in space, regardless of any small head movements the listener may make.
[0072] The reproduction of an audio scene is based on the use of acoustic transfer functions 40 of the head between the positions in space ( ^, Φ) and the auditory canal, called auditory transfer functions. These transfer functions, called "HRTF" (for "Head Related Transfer Functions"), concern the frequency form of the transfer functions. Their temporal form, i.e. the impulse response, will hereinafter be referred to as "HRIR" (for "Head Related Impulse").
[0073] In the context of binaural reproduction, the following transfer functions are then used: a frequency transfer function for the left ear (HRTFg(M)(f)), a frequency transfer function for the right ear (HRTFd(M)(f)), an impulse response function for the left ear (HRIRg(M)(t)) and an impulse response function for the right ear (HRIRd(M)(t)), M being the position of the source in polar coordinates, f the frequency and t the time.
[0074] Advantageously according to the present invention, instead of using standard transfer functions, after having modified and improved these transfer functions, reprocessed transfer functions are used, denoted in the present document HHl and HHr. HHl and HHr will be denoted their frequency form and hhl and hhr their time form. How to obtain these reprocessed transfer functions will be explained in detail later in this document.
[0075] Returning to Figure 7, the restitution of the audio scene is carried out by a binaural rendering engine. The binaural rendering engine has at its disposal the reprocessed transfer functions in their frequency form or in their time form. Thanks to these functions, for each sound source ASi of the audio scene to be reproduced, characterized by the wave temporal pattern ASi(t) of the sound source as well as its polar coordinates ( ^i,Φi) and its distance from the listener ri, the binaural rendering engine generates the corresponding three-dimensional audio scene in the two loudspeakers.
[0076] Figure 7 only represents one sound source ( ^1,Φ1,r1) but the reader will understand that what follows will be repeated for each of the sound sources, and that the respective signals of the different sound sources are added in the waves produced towards the loudspeakers (HPg, HPd).
[0077] When the binaural rendering engine works in the frequency domain, from the time pattern ASi(t) we calculate a fast Fourier transform (FFT) noted XASi(f), noted more simply XAS in figure 7. The coordinates ( ^1,Φ1,r1) of the sound source are applied on the one hand to the filter HHl for the left ear and on the other hand to the filter HHr for the right ear.
[0078] For the left channel, we multiply the fast Fourier transform XAS of the source signal by the filter HHl assigned the coordinates of the source which gives the signal to be reproduced in frequency form WASg. This signal is the subject of an inverse Fourier transform calculation 21 and we thus obtain the time signal WASg(t), i.e. image of the sound wave, which will be played on the left loudspeaker HPg.
[0079] For the right channel, we proceed in a similar manner, we multiply the fast Fourier transform of the source signal by the filter HHr assigned the coordinates of the source which gives the signal to be reproduced in frequency form WASd. This signal is the subject of an inverse Fourier transform calculation 22 and we thus obtain the time signal WASd(t), i.e. image of the sound wave, which will be played on the right loudspeaker HPd.
[0080] In an implementation when the binaural rendering engine in the time domain, for the left channel we proceed to a convolution between the signal of the sound source ASi(t) and the impulse response function hh. l(t) assigned the coordinates ( ^1,Φ1,r1) of the sound source, we thus obtain the time signal WASg(t) which will be played on the left loudspeaker HPg.
[0081] We proceed in a similar manner for the right channel, we carry out a convolution product between the signal of the sound source ASi(t) and the impulse response function hhr(t) assigned the coordinates ( ^1,Φ1,r1) of the sound source, we thus obtain the time signal WASd(t) which will be played on the right loudspeaker HPd.
[0082] It should be noted that to a certain extent the left and right transfer functions (HHl and HHr) can be personalized with respect to the listener who will experience the three-dimensional audio scene. This concerns the adaptation of the head diameter with regard to interaural delays, the azimuthal angular correction ('gamma' correction) or the combined azimuthal and elevation angular correction.
[0083] To finish on Figure 7, if the head tracking function is present, it provides in real time a signal noted HDT(t) representative of the movement of the head in yaw rotation, as well as in left-right lateral tilting and in front-back tilting, therefore three spatial rotation coordinates. As explained elsewhere, the rendering engine compensates by injecting these values with an opposite sign into the sound signal generation algorithm.
[0084] Reprocessing and Improvement of the Transfer Functions
[0085] We start from a set of primary (basic) auditory transfer functions, at the left ear and at the right ear of individuals, taken from one of N individuals, the set of primary auditory transfer functions comprising for each individual i and for each of a plurality of measured positions Mj: a fon ^. ^ ^ ^frequency transfer function for the left ear (HRTFg(i,Mj)(f)) noted ൫ , ^^ ^ ൯ and including amplitude and phase, a frequency transfer for the right ear (HRTFd(i,Mj)(f)) ^^ ൫ ^^ , amplitude and phase, a time impulse response function for the left ear (HRIRg(i,Mj)(t)) denoted ℎ ^ ൫ ^^ ^ , ^^ ^ ൯ and a time impulse response function for the ear (HRIRd(i,Mj)(t)) denoted ௗ ^ e right ℎ ൫ ^^ ^ , ^^ ^൯, each position Mj having the direction in polar coordinates ^j and Φj,
[0086] The frequency transfer functions are in this document manipulated in the digitized / digital form, namely a sequence of values. For a large number of frequencies located in the human audible spectrum, the transfer function gives an amplitude and a phase. There can typically be 128 frequency points listed or even 256 frequency points listed. For each of the amplitude and the phase are coded in digital format on 8, 10 or 12 bits.
[0087] The different stages of reprocessing of these raw transfer functions are illustrated in Figure 8. The first stage noted S0 corresponds to the provision of the raw transfer functions mentioned above from measurements with, where appropriate, right-left symmetrization.
[0088] It is expected that the step of symmetrization of the spectral amplitude SA in addition a smoothing by local average of k neighboring directions ^^, the formulation: ^^ ^^. ^ 1 = ^ ^ ^^ ^^ ൬ ^^ ^^ ^^ ^ ^ ^ ^ ^ ^^ ^^ ^^ ^^ + ^^ ^^ ^ ^^ ^^ ^^ ^ ^ ^ ^ ^^ ^^ ^^
[0089] The this step in addition to the passage in logarithm, if necessary symmetrization of spectral for all directions ( ^^, ^^) of a subject ^^, using an averaging function FA, according to the formulation: ^^ ^ ( ^^, ^^) = ^^ ^^ ^ ^^ ^^ ( ^^ ^ ^ ^ ^ ( ^^ ^^ , ^^ ^^ )) , ^^ ^^ ( ^^ ^ ^ ^ (− ^^ ^^ , ^^ ^^ ))^ if the digitized, then the SA step is logarithmic.
[0091] According to a simple example, the averaging function FA consists of the half-sum of the two values.
[0092] Then we proceed to a step noted SB, of smoothing the spectral amplitude by binaural relative deviation, normalization for a reference direction ( ^^ ^ , ^^ ^ ) with average over the subjects, according to the formulation, for a source on the left: ே ^^ ^^ ൫ ^^ ^ ^^, ൯ = 1 ^ ^ ^ ^ ^^ ^ ^^ − ^^ ^ ൫ ^^0, ^^ ൯^
[0093] The and Φ0=0. It of According to another solution, the reference direction ^0,Φ0 can correspond to ^0 =30° and Φ0=0. These are so-called 'audio' filters in the context of the present invention.
[0095] For a right-hand source, the normalization can be applied in a similar manner with respect to the reference direction ^0,Φ0 on the right side: ^^ ^^ ൫ ^^ ^ ( ^^, ^^) ൯ = 1 ^ ^ ^ ^ ^^ ^( ^^) − ^^ ^ ൫ ^^0, ^^ ൯^ ൯
[0096] noted SB2 CFg according to: ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = ^^ ^^ ^ ^^ ^ ൫ ^^0, ^^0൯^ for any direction ^j , Φj included in the angle 0, +Φ0] side can correspond to terminals ^0 = 30° and Φ0 = 12°. According to a non-limiting example, the front cap CFg can correspond to terminals ^0 = 20° and Φ0 = 8°. It is understood that the extent of the front cap CF can be in view of the intended applications, eg single source audio scene multi-source audio scene.
[0100] The above formula is given for a source located on the left. If the source is located on the right then ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = ^^ ^^ ^ ^^ ^ ൫ ^^0, ^^0൯^ for any direction ^ j , Φ j included in the solid angle CFd defined by the respective terminals 0, and [- Φ 0, +Φ0] right side.
[0101] Said the frontal cap CFg or CFd is specific to the side of the and confined to the side of the source, i.e. it is limited by the boundary CC ^ = 0°.
[0102] The flattening of the filters on a frontal cap (CF), can be carried out in the formulation otherwise written: ^^ ^^൫ ^^ ^ ^^ ൯ = 1 ே ^ ^ ^ ^ ^^ ^ ^^, ^^ − ^^ ^ ^^ , ^^ ^ ൯
[0103] He is the source, not just the one for a source located on the right.
[0104] It should also be noted that the cap CFg or CFd is not necessarily square or rectangular, it can be circular or elliptical, for example it can be delimited by the equation ^² + exc Φ² < K, inside the hemisphere of the source, exc being the eccentricity of the ellipse, K being a parameter, this variant is reference CF' in figure 6.
[0105] Furthermore, we proceed to a step noted SC of symmetrization and average of the phase, according to the ே formulation: < ^^ ^ ( ^^, ^^) > = 1 2 ^^ ^ ^ < ^^ ^ ^ ( ^^, ^^) > + < ^^ ௗ ^ ^(− ^^ ^^ , ^^ ^^ ^ > ^ >
[0106] The notation < . > denotes the of the complex transfer function. The SC step can be carried out at any time relative to the other offline steps, notably for example at the very beginning in parallel with the SA step later.
[0108] Furthermore, we proceed to a step denoted SD, for evaluating the average delay; ே ^^ ^ = 1 ^ ^ ^ ^^ ^^ ^^ ^^ ^^ ^^൫หℎ ^ ^ ∗ ^^ห൯ fs being the HRIR, and ^^ ^^ ^^ ^^ ^^ ^^( ^^ ^ ) being , digitized values of the impulse response, the notation * being convolution product, w being a convolution filtering function. The convolution filtering function w can be reduced to a Dirac impulse (then we simply select the largest signal peak).
[0109] Note that the evaluation of the average interaural delay can be calculated using other methods, without necessarily using the formulas above.
[0110] The SC and SD steps can be carried out in the order in which they are here, but the order on this point can be different, the SC and SD steps being carried out before or after the other steps.
[0111] We can also carry out a step denoted SE of interpolation of the spectral amplitudes, phases and interaural delays. This interpolation can consist of a linear interpolation between the known points. According to another example, it is a second-degree polynomial interpolation between the known points.The final granularity can consist of providing a value every 3°, or even a value every 2°, in azimuth and elevation. As will be seen later, this SE interpolation step can be carried out in offline calculations or in the real-time process.
[0112] Then we proceed to a step noted SF, of generating a set of reprocessed transfer functions, noted HHl and HHr, ^^ ^^. ^ ( ^^, ^^) = 10ௗ^൫ு^(ఏ,థ)൯ / ଶ^
[0113] This step from the mode If we take the right ear as the first ear, instead of the left ear as above, then this SF step is written as follows: ^^ ^^ ^ ( ^^, ^^) 10ௗ^൫ு^(ఏ,థ)൯ / ଶ^
[0115] But in fact this comes down to the change of sign for angular phase shift. We note that the generation of a set of reprocessed transfer functions could be calculated differently, without necessarily using formulas 5 above.
[0117] Then we can proceed to a step of individualization of the interaural delay SG, as a function of a ^ head radius b of a target listener, the delay being determined by ^^^ ^^, with b the new head radius and a the original radius. 10
[0118] The radius a can come from the measurements accompanying the transfer function measurements. The radius a can come from anthropological data.
[0119] The radius b can be chosen according to the listener benefiting from the audio scene rendering. For example, the radius b can be smaller than a for someone with a small head circumference, we will speak of size XS or S. 15
[0120] For example, the radius b can be larger than the radius a for someone with a large head circumference, we will speak of size L or XL.
[0121] For example, the radius b can be larger than the radius a for someone with an average head circumference. Then we proceed to a step denoted SH of angular correction.20 According to a first simple example called here 'in two dimensions' we make an azimuthal angular, with a parameter γ which can be adapted to the user, according to the formulation: ^^( ^^, ^^) = ( ^^ + ^^ ∗ sin(2 ^^), ^^)
[0124] In other words, ^^ directly, we use instead a. form ^^ + ^^ ∗ sin(2 ^^). The value of γ is configurable. If the 'head tracking' function is implemented, the SH step to the yaw motion corrections on ^ in a similar way to that presented above. 30
[0127] According to another example of the SH step of angular correction, we correct in azimuth and elevation, according to a so-called 'three-dimensional' deformation, we use f( ^^, ^^) = ( ^^ + ^^ ∗ sin(2 ^^), ^^ + ^^ ᇱ∗ sin( ^^)). The parameters γ and γ' can be adapted to each user.
[0128] In an even more generic way, the step noted SH of angular correction 35 consists in replacing the angles θ,φ in the HRTF formulas by the angles (θ′, ^^′) such that (θ′, ^^′) = f (θ, ^^) f being any function of correction in azimuth and elevation.
[0129] Figures 8, 9 and 10 schematically illustrate the steps described above of the proposed method, according to three variants of hardware implementation. 40 functional distribution differs between steps carried out in offline mode and steps carried out in real-time mode.
[0130] In Figure 8, only the steps SH and SR are carried out in the target device, for example an audio headset, in real time compared to the execution of the binaural rendering. The other steps were carried out upstream, by calculation, in one or more computers. We note in particular that the SE interpolation step was carried out upstream.It should also be noted that, as mentioned previously, certain steps are optional, for example steps SC, 5 SG. It is understood that if step SG is not performed, the rendering is not personalized according to the head radius of the target listener. It is noted that the azimuthal angular correction step SH (and where appropriate in elevation) can be personalized according to the target listener.
[0131] In Figure 9, steps SG, SH and SR are performed in the target device 10, for example an audio headset. Here, step SG can then be personalized according to the target listener, just like step SH. Steps SA to SF are performed upstream, by calculation, in one or more computers.
[0132] In Figure 10, steps SE, SF, SG, SH and SR are performed in the target device, for example an audio headset.In particular, the SE interpolation step 15 can be performed in real time for the positions that are actually necessary and used by the binaural rendering engine. This makes it possible to have a memory space of limited size and to calculate the interesting points at the last moment.
[0133] The SG and SH steps can be customized according to the target listener 20. Only the SA to SD steps are performed upstream, by calculation, in one or more computers.
[0134] The choice between the different solutions can be dictated by the audio production chain or by the different hardware limitations in terms of computing capacity or in terms of available memory size. 25
[0135] Figures 2 and 3 illustrate filters normalized with respect to the reference direction ^. 0= 15°, this normalization being produced by the SB step described above.
[0136] In Figure 2, curve 61 (solid line) represents the amplitude of the frequency transfer function, for the contralateral ear, for ^ = 0°, 30 normalized with respect to the reference direction ^ 0= 15°.
[0137] Curve 62 (small dotted line) represents the amplitude of the frequency function for ^ = 15°. Curve 63 (dash line) represents the frequency transfer function for ^ = 30°. Curve 64 (long dashed line) represents the amplitude of the frequency transfer function for ^ = 45°.
[0138] Curve 65 (solid line) represents the phase of the frequency transfer function, for the contralateral ear, for ^ = 0°, normalized with respect to the reference direction ^0 = 15°. Curve 66 (small dotted line) represents the phase of the frequency transfer function for ^ = 15°. Curve 67 (dash line) represents the phase of the frequency transfer function for ^ = 30°. Curve 68 (long dotted line) represents the phase of the frequency transfer function for ^ = 45°.
[0139] In Figure 3, the curve 71 (continuous) represents the amplitude of the frequency transfer, for the ipsilateral ear, for ^ = 0°, relative to the reference direction ^. 0= 15°.
[0140] Curve 72 (small dotted line) represents the amplitude of the frequency transfer function for ^ = 15°. Curve 73 (dash line) represents the amplitude of the frequency transfer function for ^ = 30°. Curve 74 (solid line) represents the amplitude of the frequency transfer function for ^ = 45°.
[0141] Curve 75 (solid line) represents the phase of the frequency transfer function, for the ipsilateral ear, for ^ = 0°, normalized with respect to the reference direction ^0 = 15°. Curve 76 (small dotted line) represents the phase of the frequency transfer function for ^ = 15°. Curve 77 (dash line) represents the phase of the frequency transfer function for ^ = 30°. Curve 78 represents the phase of the frequency transfer function for ^ = 45°.
[0142] We note that curve 72 is essentially flat, which confirms the normalization with respect to ^=15° for the ipsilateral ear.Figures 4 and 5 illustrate filters normalized with respect to the reference direction ^0 = 30°, and flattened in a frontal cap area, these transformations being produced by the steps SB and SB2 described above.
[0144] In Figure 4, curve 81 (solid line) represents the amplitude of the frequency transfer function, for the contralateral ear, for ^ = 0°, normalized with respect to the reference direction ^0 = 30°.
[0145] Curve 82 (small dotted line) represents the amplitude of the frequency transfer function for ^ = 15°. Curve 83 (dash line) represents the amplitude of the frequency transfer function for ^ = 30°. Curve 84 (solid line) represents the amplitude of the frequency transfer function for ^ = 45°.
[0146] Curve 85 represents the phase of the frequency transfer function, the contralateral ear, for ^ = 0°, normalized with respect to the direction of ^0 = 30°.Curve 86 represents the phase of the frequency transfer function for ^ = 15°. Curve 87 represents the phase of the frequency transfer function for ^ = 30°. Curve 88 represents the phase of the frequency transfer function for ^ = 45°.
[0147] In Figure 5, curve 91 (solid line) represents the amplitude of the frequency transfer function, for the ipsilateral ear, for ^ = 0°, relative to the reference direction ^0 = 30°, with flattening of the amplitude from 0° to ^0 = 30.
[0148] Curve 92 (small dotted line) represents the amplitude of the frequency function for ^ = 15°. Curve 93 (dash line) represents the amplitude of the frequency transfer function for ^ = 30°. Curve 94 (solid line) represents the amplitude of the frequency transfer function for ^ = 45°.
[0149] Curve 95 (solid line) the phase of the transfer function for the ipsilateral ear, for ^ = 0°, normalized with respect to the reference ^.0= 30°. Curve 96 (small dotted line) represents the phase of the frequency transfer function for ^ = 15°. Curve 97 (dotted line) represents the phase of the frequency transfer function for ^ = 30°. Curve 98 (solid line) represents the phase of the frequency transfer function for ^ = 45°.
[0150] Note that curves 91,92,93 are flat, confirming the effect of flattening and normalization to 30°, over the entire range from 0° to 30° (0°, 15° and 30° identical here).
[0151] It can be seen from these curves that the contralateral ear is affected in the same way in the HRTFs normalized to 15° ('image') and normalized to 30°; considerations 15 reprocessed filters produced by steps SA to SG or SH constitute an improved data set that can be made available to third parties for a license of use or as part of a sale.
[0154] The inventors have noticed that other competing bases such as those of Dolby™ or similar have more choppy amplitudes between 1kHz and 10kHz and give less good results.
[0155] Optionally, due to the symmetrization, it is possible to keep in memory the filters only for one hemisphere, right or left, and to use a plane symmetry with respect to the sagittal plane PMS carried out by the real-time binaural rendering engine.
[0156] More precisely, in practice, after the symmetrization steps of the phase and the amplitudes SA and SC, it is possible to keep data only for one hemisphere and make the complementary transformations at steps SB, SB2, SD only on one hemisphere. The process for the other hemisphere will be reconstructed in real time by plane symmetry with respect to the sagittal plane 30 PMS.In this way, it is possible to occupy a memory space divided by 2, which may be relevant for certain categories of hardware or audio processing chains.
[0157] According to another embodiment of the present invention, starting from certain HRTF filter bases, possibly already symmetrized, the flattening defined in step SB2 explained above is simply applied. In this case, it may be necessary to connect by continuity the amplitudes of the filters to the edge of the cap CFg, (respectively CFd) to guarantee the absence of jumps. This can be done by a polynomial interpolation.
[0158] According to another optional characteristic, in addition to the transformations 40 above, an additional transformation may be provided to raise the high-pitched sounds collectively on all the filters.Thus, for example, a reinforcement of sounds with a frequency higher than 2 or 3 kilohertz can be practiced, this reinforcement being of an amplitude between 1 dB and 3 dB for example.
[0159] This is illustrated on the where curves 91, 92, 93 show an increase in amplitude above 3 kilohertz. It is noted that because of this characteristic, the low frequency plateau has been shown shifted, below 0 dB. 5.
Claims
CLAIMS 1. Method for generating an audio scene in a binaural spatialization system, the method being characterized in that it comprises: S0- a step of providing a set of primary auditory transfer functions, relating to the left ear and the right ear of individuals, taken from a set of N individuals, the set of primary auditory transfer functions comprising for each individual i and for each of a plurality of measured positions Mj: a frequency transfer function for the left ear (HRTFg(i,Mj ^ ^ )(f)) noted ^^ ൫ ^^ ^ , ^^ ^ ൯ and including amplitude and phase, a frequency transfer function for the right ear (HRTFd(i,Mj)(f)) noted ൫ and including amplitude and phase, a response function (HRIRg(i,Mj)(t)) noted ^ ^ temporal impulse nse for the left ear ℎ ൫ ^^ ^ , ^^ ^൯ and a time impulse response function for the right ear (HRIRd(i,Mj)(t)) denoted ℎ ௗ ^ e ൫ ^^ ^ , ^^ ^൯ , each position Mj having as direction in polar coordinates ^j and Φj, the process comprising: SA- a preliminary preparation step noted SA, in which a passage in logarithmic values is carried out, with where appropriate symmetrization of the spectral amplitude for all directions ( ^^, ^^) of a subject ^^, using an averaging function FA, according to the formulation: ^^ ^ ( ^^, ^^) = ^^ ^^ ^ ^^ ^^ ( ^^ ^ ^ ^ ^ ( ^^ ^^ , ^^ ^^ )) , ^^ ^^ ( ^^ ^ ^ ^ ^ (− ^^ ^^ , ^^ ^^ )) ^ ^ SB- a by binaural relative deviation, normalization for a reference direction ( ^^ ^ , ^^ ^ ) with average over the subjects, according to the formula: ^ ^ ^^൫ ^^^^^ ൯ = 1 ^ ^ ^ ^ ^^ ^ ^^ − ^^ ^ ൫ ^^ 0 ^^ ൯^ ൯ SD- a noted ^^( ^^, ^^); SF- a step, noted SF, of generating a set of reprocessed transfer functions, noted respectively ^^ ^^^ ( ^^, ^^ )^^ ^^ ^^ ^^^( ^^, ^^ ) SR- a step noted SR of using the reprocessed functions ^^ ^^ ^ ( ^^, ^^) and ^^ ^^ ^ ( ^^, ^^) to generate an audio scene in real time using a binaural rendering engine, with one or more point audio sources (s) prescribed.
2. Method according to claim 1, in which in step SD-, the average intraural delay noted ^^( ^^, ^^) is calculated as follows: ^^ ^ ^^ = 1 ^ ^ ^ ^^ ^^ ^^ ^^ ^^ ^^൫หℎ ^ ^ ^ ∗ ^^ห൯ fs being the ^^ ^^ ^^ ^^( ^^ ^ ) being the index of the maximum of a w being a convolution filtering function.
3. Method according to claim 2, wherein in step SF-, the set of reprocessed transfer functions is calculated as follows: ^^ ^^ ^ ( ^^, ^^) = 10ௗ^൫ு^(ఏ,థ)൯ / ଶ^ ) 4. Method according to the noted SB2 of flattening of the amplitude of the filters on a frontal cap (CF) according to the formulation: ^^ ^^ ൫ ^^ ^ ^^ ൯ = 1 ே ^ ^ ^ ^ ^^ ^ ^^ − ^^ ^ ^^ ^ for all the limits [- 0, - Φ0] and [ ^0, 5. Method according to claim 3, further comprising a step denoted SB2 of flattening the amplitude of the filters on a frontal cap (CF), with ^^ ^^൫ ^^ ^ ( ^^, ^^)൯ = ^^ ^^ ^ ^^ ^ ൫ ^^0, ^^0൯^ for any direction ^j , Φj included in the angle 0 of the left side and ൯ ൫ , included in solid CFd respective [0, ^ ] and 0, +Φ0] of the side 6. Method according to any one of claims 1 to 5, wherein the binaural rendering engine further comprises a listener head movement tracking function.
7. Method according to any one of claims 1 to 6, wherein it is further provided, before step SF or during step SR: a step denoted SE of interpolation of spectral amplitudes, phases and interaural delays.
8. Method according to any one of claims 1 to 7, in which the step of symmetrization of the spectral amplitude (SA) further comprises a smoothing by local average of k neighboring directions, according to the formulation: ^^ ^ 1 = ^ ^^ ^^ ൬ ^^ ^^ ^^ ^ ^ ^ ^ ^ ^^ ^^ ^^ ^^ + ^^ ^^ ^ ^^ ^^ ^^ ^ ^ ^ ^ ^^ ^^ ^^ 9. further a step of individualizing the interaural delay (SG) as a function of a head radius b of a target listener, the interaural delay being determined by ^ ^^, with ^^ the new head radius and ^^ the original radius.
10. Method according to any one of claims 1 to 9, further comprising a step denoted SH of angular correction, with a parameter γ which can be adapted to each user, according to the formulation: ^^( ^^, ^^) = ( ^^ + ^^ ∗ sin(2 ^^), ^^) 11. Method according to any one of claims 1 to 10, further comprising a step denoted SC of symmetrization and phase average, according to the formulation: ே < ^^ ^ ( ^^, ^^) > ^^ ^ ^ < ^^ ^ ^ ( ^^, ^^) > + < ^^ ௗ ^ ^ (− ^^ ^^ , ^^ ^^^ > ^ ^ the notation 12. Audio headset configured to implement steps SE to SR of a method according to any one of claims 1 to 11.
13. Processing system configured to implement steps S0 to SG of a method according to any one of claims 1 to 11.