Interaural sound reproduction using multi-order HRTFs
Multi-order HRTFs address the challenge of stable sound localization by encoding sound position across multiple orders, enabling accurate virtual surround sound reproduction using stereo speakers, suitable for various audio applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2026-03-12
AI Technical Summary
Current surround sound technologies, including headphones and home setups, struggle to accurately reproduce a sound field outside the listener's head, with existing HRTF methods failing to provide stable sound localization due to individual variations and latency issues.
The use of multi-order head-related transfer functions (HRTFs) for sound localization, involving multiple orders of HRTFs to encode sound position, allowing for inter-individual averaging and stable localization of sound sources outside the listener's head, and embedding positional information in stereo speaker sound to create virtual surround sound fields.
Multi-order HRTFs enable accurate sound localization and creation of virtual surround sound fields using two stereo speakers, suitable for applications requiring low latency, such as VR and gaming, and applicable to both headphones and speakers.
Smart Images

Figure 0007828960000001 
Figure 0007828960000002 
Figure 0007828960000003
Abstract
Description
[Background technology]
[0001] Increasing listener involvement and immersion in recorded and reproduced sound has long been a goal in the audio industry. This quest began as early as 1931, when Alan Blumlein invented stereo. Since then, sound quality and immersion have steadily improved. While surround sound has existed in various forms for a long time, it wasn't until the 1970s that Dolby introduced Dolby Stereo, which, despite its name, became the first commercially successful surround sound format. Surround sound offered a previously unattainable level of immersion. In recent years, object-based audio formats like Dolby Atmos and Sony 360 have emerged, further enhancing immersion.
[0002] One of the major challenges common to all surround formats is reproducing the surround sound field. While commercial Dolby Atmos theaters, with hundreds of speakers positioned around the perimeter of a room, can deliver impressive sound, recreating such a setup in a private home is impractical. The industry also struggles with reproducing the surround sound field with headphones. Despite significant research efforts, current technology is unable to reproduce a sound field far outside the listener's head. Sound typically feels mostly inside the head, rather than surrounding the listener as intended. Furthermore, the small amount of sound outside the listener's head is typically located immediately to the left or right of the listener's ears or slightly behind them. This clearly makes it impossible to provide the highly desirable, stable positioning of the front hemisphere.
[0003] To locate sound sources in space, so-called head-related transfer functions (HRTFs) are commonly applied. Surround sound produced for movies and games, as well as many stereo recordings, include HRTF encoding of sound. HRTF position encoding is present in both surround sound and stereo recordings and is suitable for both loudspeaker and headphone playback. Some playback algorithms, such as Dolby Atmos for headphones, also employ HRTF encoding to localize sound.
[0004] Several HRTF databases containing measurements from hundreds of subjects have been published on the web by the research community and are available for download. Databases typically contain frequency responses associated with multiple locations around each subject, known as head-related frequency responses (HRFRs), and in some cases, associated time-domain responses called head-related impulse responses (HRIRs).
[0005] Typically, the HRFRs for several hundred individuals are averaged to generate an average HRFR for each location. The average HRFR data is then used to encode the location of sound sources during recording and playback.
[0006] As mentioned above, this type of HRFR coding does not give convincing results for headphones, requires multiple speakers placed around the room, and the perceived position varies greatly from person to person, despite averaging measurements from multiple subjects.
[0007] However, good results can be obtained by measuring the HRIRs of each listener individually. Convolving the playback material with the individual HRIRs using a conventional FIR filter can achieve fully realistic immersion in surround sound with headphones, but this is only possible for people whose individual HRIRs are used in the playback convolution. It is obviously impossible to create individual HRIR data for every person who will listen to a recording. Several attempts have been made to customize commonly used average HRFR data from information provided by individuals about their physical characteristics, but none have been successful.
[0008] HRIR also has the problem of filter latency: it needs to be fairly long to get good results, but the latency introduced poses a big problem in virtual reality, gaming, and similar applications that cannot tolerate large latency.
[0009] Simple averaging approaches like HRFR averaging are also unsuccessful in the time domain. Figure 1 illustrates the challenges of time-domain HRIR averaging. Traces 1, 2, and 3 in Figure 1 are HRIR data from three different subjects. Due to differences in body size and associated sound wave travel times, the second step in the HRIR data occurs at a different time than the large initial arrival on the left side of the trace. Trace 4 represents the average of traces 1, 2, and 3. This is clearly not an appropriate average for three physically different subjects. In this example, trace 2 is the best average for the individual body sizes, but trace 4 is completely different from trace 2. The three steps in traces 1–3 are blurred in time. Instead of a clear wavefront reaching the temporal average point of trace 2, the wavefront is blurred and suppressed in time, which is not the desired result.
[0010] This invention solves position coding by decomposing the localization process in a novel way, introducing a new approach focused on the time domain. This approach is called multi-order HRTFs. This approach allows for inter-individual averaging, and time-domain coding provides more stable localization of sound sources clearly located outside the listener's head, or in front of them if desired, through headphones. Additionally, by embedding coded position information in the direct sound from a pair of stereo speakers, virtual surround sound sources can be created around the listening room using only two stereo speakers. Summary of the Invention
[0011] The present invention relates to a method for sound reproduction, comprising positional coding with multi-order head-related transfer functions (HRTFs), where sound is reproduced with at least a first-order HRTF to the left ear, then a second-order HRTF from the left ear to the right ear, and simultaneously a first-order HRTF to the right ear, then a second-order HRTF from the right ear to the left ear. In the above context, "multiple orders" means second, third, or any order up to any level. In this regard, according to one embodiment, the method may include at least a third-order HRTF from the left ear to the right ear as well as from the right ear to the left ear, and preferably at least a fourth-order HRTF from the left ear to the right ear as well as from the right ear to the left ear. [Brief explanation of the drawings]
[0012] The concepts of the present disclosure are also described below with reference to the drawings, and in particular with reference to FIG.
[0013] Furthermore, in relation to the present invention, there are many known methods using several / multiple HRTFs, such as those disclosed in US2020 / 0037097, but these are not the same concept as that disclosed and provided by the present invention. Again, the present invention provides a method that includes sound reproduction using at least a primary HRTF to the left ear, then a secondary HRTF from the left ear to the right ear, and simultaneously a primary HRTF to the right ear, then a secondary HRTF from the right ear to the left ear. This should not be confused with the use of several / multiple HRTFs utilized in many known methods.
[0014] Detailed explanation of multi-order HRTFs It is well known from psychoacoustic research that human hearing is extremely sensitive to the time domain characteristics of sound. The difference between the sound of wood and metal can be heard within the first few milliseconds after striking the material. The attack waveforms of violin and trumpet sounds are very different, and the difference is easily audible. However, when listening to sustained sounds from each instrument without an attack, it is difficult to distinguish between the two.
[0015] Similarly, the location of a sound source can be interpreted not only from HRFR but also from time-domain information. Due to these difficulties, traditional localization tasks have focused on average HRFR data, ignoring time-domain information. However, the results have been poor. Personal HRFR data captures time-domain information, but only for one individual at a time, and is successful in providing a good impression of the surround sound field for that individual.
[0016] Figure 2 shows the sound path from a sound source to the listener's head and its surroundings. 1 is the listener, 2 is the sound source, and 3-8 visualize the sound wave paths to the head and its surroundings. Figure 2 shows the location of one sound source, but a similar sound path can be assumed for any position in three-dimensional space. Figure 2 shows the general principle, and paths for other sound source positions can be easily estimated.
[0017] Each sound path (3-8) has an associated time delay, frequency response, and attenuation. Path 3 has a time delay, which is the sound's travel time from sound source 2 to the right ear, but this particular case has zero delay because it is the first arrival of the sound at the listener and does not require a parallel delay in the sound's travel time to reach the listener. The attenuation of this particular first-order path is also zero because the sound reaches the ear directly without any obstructions to cause attenuation. The frequency response is typically the well-known average HRFR for the sound source location at the right ear. However, the sound wave does not stop upon reaching the right ear. The sound wave travels around the head along path 6 to the left ear. This path has an interaural time delay due to sound travel time, a frequency response due to high-frequency shadowing by the head, and attenuation due to traveling around the head to the opposite ear. This second wave path is the second-order HRTF. After the sound wave reaches the left ear, it returns to the right ear via path 8, again with associated time delay, frequency response, and attenuation. This is the third-order HRTF. For clarity, Figure 2 does not show higher order HRTFs, but the principle is clear and higher order HRTFs can be easily estimated by simply following the path around the head.
[0018] The time delay associated with the interaural path is directly related to the physical distance between the ears and ranges from 200 μs to 1 ms, with a typical value of approximately 600 μs. The frequency response change caused by the head as sound waves travel from one ear to the other is generally a downshelving of the high-frequency spectrum beginning at 400 Hz to 2.5 kHz and continuing up to 20 kHz, the limit of human hearing. Additionally, due to the physical characteristics of the human head and shoulders, there are several dips and peaks associated with specific paths. Attenuation typically varies from 0 to 6 dB for the primary path, 3 to 12 dB for the secondary path, 6 to 24 dB for the tertiary path, and 9 to 48 dB for the quaternary path. Methods and techniques for obtaining the precise time delay and attenuation associated with each path are straightforward for those skilled in the art using standard methods and will not be described here.
[0019] The relevant frequency response can be determined from readily available HRTF data. Figure 3 shows the frequency response associated with sound location 2, sound path 6 in Figure 2, as magnitude (dB) versus frequency (Hz).
[0020] Acoustic measurements, as mentioned above, show that sound waves propagate around an object several times, and that adding second-, third-, and fourth-order HRTFs makes the sound feel more natural and significantly improves sound source localization. Localization and naturalness improve with each additional order up to fourth order, but become less noticeable beyond that. Of course, HRTFs can be used up to any possible order, from second order to hundreds or even thousands, but as mentioned above, it is understood that only minor benefits are obtained beyond fourth order.
[0021] Similarly, the sound paths from the sound source to the left ear, starting from path 4, each have a time delay, frequency response, and attenuation similar to the path starting from path 3 above. However, the time delay of path 4 is not zero, as with path 3, but is due to the time difference between the two ears. The resulting frequency change is typically the known average HRFR for the sound source location at the left ear. The attenuation along path 4 is typically 4.5 dB when the sound source is positioned as shown in the example. The next second-order and third-order paths 5, and the subsequent path 7, are also associated with time delays, frequency responses, and attenuation.
[0022] The sound path from one ear to the other through the front of the head is slightly longer than the sound path through the back of the head. This sound path also experiences slightly different attenuation and frequency changes than the sound path through the back of the head. Taking this into account, the head and ears are excellent localization devices, capable of generating a series of unique multi-order HRTF sound paths for different sound source positions. As a result, multi-order HRTFs can achieve stable localization of sound sources both in front of and behind the head.
[0023] Because multi-order HRTFs separate the variations in frequency response, the attenuation and time delay of each path averaged across subjects is not complicated. The frequency response of multiple individual paths can be easily averaged using existing methods, and the attenuation and delay are simply the average of the attenuation and travel distance of each path for each subject. Averaging the characteristics of multiple individuals is important to obtain consistent, similar results for all listeners.
[0024] The frequency changes associated with each path can be easily implemented using standard IIR filters, eliminating the latency associated with FIR filters. Therefore, multi-order HRTFs operate without any latency and are suitable for applications such as VR, gaming, and applications requiring zero or very low latency. Figure 4 is a block diagram of a typical DSP implementation of multi-order HRTFs. A fourth-order implementation for one sound source position is shown. Of course, multi-order HRTFs can be implemented in many other ways, and it should be clear that Figure 4 illustrates only one example of many possible topologies. Blocks 11, 21, 31, 41, 51, 61, 71, and 81 are delay blocks that apply the delays associated with each set of four paths for each ear in a fourth-order implementation. Blocks 12, 22, 32, 42, 52, 62, 72, and 82 apply the frequency changes associated with each path. Blocks 13, 23, 33, 43, 53, 63, 73, and 83 are gain blocks that apply the attenuation present in each path. Finally, 100 is a summer block that simply sums all the outputs from the four paths to the left ear, and 200 is the summer for the right ear. The outputs from 100 and 200 are sent to the left and right channels, respectively.
[0025] Applications that utilize multi-order HRTFs may have both stereo and multi-channel input signals. Multi-order HRTFs allow for the creation of multiple virtual sound sources. If the input signal is in a conventional five-channel surround sound format, the multi-order HRTFs can be used to create five virtual speakers in the usual locations of a five-channel surround sound setup, e.g., front left and right, center, and surround left and right. Each individual input channel is then played back through the corresponding virtual speaker. Similarly, with more surround speakers and additional ceiling speakers, more virtual speakers can be created. For stereo input signals, the individual feeds can be extracted to the virtual speakers using conventional sound extraction and steering processes. In this case, the stereo extraction and steering processes are the same as for conventional surround sound products.
[0026] Virtual sound sources created with multi-order HRTFs work with both headphones and speakers. For headphones, it is possible to create a surround sound field that is close to the experience of the human body using personally measured HRIRs. For speakers, it is possible to encode the virtual speakers into the sound from a pair of stereo speakers that create a virtual center speaker, surround speakers, and height speakers. With multi-order HRTF virtual speakers, it is possible to create a surround sound field similar to that achieved by installing multiple speakers.
[0027] Reproduction using multi-order HRTF virtual sources is, of course, not limited to current stereo and surround formats and the locations of those sources. The above examples are merely illustrative of possible multi-order HRTF applications, and any desired number of virtual speakers may be created in any location.
[0028] Multi-order HRTFs can be applied at any stage from sound recording / generation to playback, and are not limited to the playback stage. Multi-order HRTFs can be used in design and / or production to apply positioning to sounds that will be played over headphones, conventional stereo, or multi-channel playback systems. For example, multi-order HRTFs can be used in game engines to localize sounds within the game's generated sound field. As another example, multi-order HRTFs can be used within DAW software as an integration or plug-in to localize sounds within the sound field in sound production. In other words, multi-order HRTF algorithms and acoustic processing can be applied at any stage and achieve the same results.
[0029] Below are some specific embodiments of the present disclosure.
[0030] According to one particular embodiment of the present disclosure, the method includes at least a third order HRTF going from left ear to right ear as well as from right ear to left ear, and preferably at least a fourth order HRTF going from left ear to right ear as well as from right ear to left ear.
[0031] Additionally, according to another embodiment, the method includes creating one or more virtual sound sources by embedding encoded location information in the sound.
[0032] According to yet another embodiment, each head-related transfer function (HRTF) of second or higher order includes parameters such as time delay, frequency response, and attenuation.
[0033] Furthermore, according to another embodiment, the method considers the difference between different sound paths, for example the difference between a sound path from one ear to the other in the front of the head and a sound path at the back of the head, where the sound path from one ear to the other is any path around the head. Thus, the method according to the present disclosure may include multiple sound paths.
[0034] Also, according to yet another embodiment, the method includes averaging. As mentioned above, according to the present disclosure, inter-individual averaging is possible. In the case of time domain encoding, this provides a more stable localization of sound sources clearly located outside the listener's head, and in front of them if desired. Based on this, according to one embodiment of the present disclosure, the method includes averaging focused on the time domain. Furthermore, according to one embodiment of the present disclosure, the method includes averaging of parameters such as time delay, frequency response, and attenuation that are independent of each other. This is a further difference when compared to averaging by known methods used today.
[0035] Additionally, the present disclosure is directed to different types of systems, hardware and software implementations.
[0036] According to one embodiment, the present disclosure is directed to a headphone playback system configured to use the method according to the present disclosure.
[0037] Additionally, the present disclosure is directed to a speaker playback system configured to use the methods of the present disclosure.
[0038] Furthermore, the present disclosure is directed to a playback system comprising a pair of stereo speakers, the system being configured to use the method according to the present disclosure to create virtual surround sound sources around a listening room by embedding encoded positional information in the direct sound from the pair of stereo speakers.
[0039] As is apparent from the above, other applications are possible with the present disclosure.
[0040] According to one embodiment, the present disclosure is directed to a gaming engine system configured to use the methods of the present disclosure. According to another embodiment, the present disclosure provides a digital audio workstation (DAW) software system configured to use the methods of the present disclosure.
Claims
1. Position coding with multi-order head-related transfer functions (HRTFs), Reproducing sound at least through a first order HRTF and a first sound path to the left ear, then a second order HRTF and a second sound path from the left ear to the right ear, and simultaneously a first order HRTF and a first sound path to the right ear, then a second order HRTF and a second sound path from the right ear to the left ear; including at least third order HRTFs from the left ear to the right ear as well as from the right ear to the left ear, and third order sound paths from the left ear to the right ear and from the right ear to the left ear, respectively; Taking into account the differences between different sound paths, How to play sound.
2. At least a fourth-order HRTF going from the left ear to the right ear as well as from the right ear to the left ear; 2. The sound reproduction method of claim 1, comprising:
3. 3. A method of reproducing sound according to claim 1 or 2, comprising creating one or more virtual sound sources by embedding encoded position information in the sound.
4. 4. A sound reproduction method according to claim 1, wherein each of the head-related transfer functions (HRTFs) of second or higher order includes parameters for time delay, frequency response and attenuation.
5. A sound reproduction method according to any one of claims 1 to 4, wherein the difference is the sound path from one ear to the other at the front of the head and the sound path at the back of the head.
6. A sound reproduction method as claimed in any one of claims 1 to 5, focusing on the time domain and comprising averaging of independent time delay, frequency response and attenuation parameters between subjects.
7. A headphone playback system comprising headphones and a DSP including blocks that apply at least the first-order HRTF, the second-order HRTF and the third-order HRTF, and an adder block that sums the outputs of at least the first-order sound path, the second-order sound path and the third-order sound path for each of the left and right ears, and which creates multiple virtual sound sources using the method of any one of claims 1 to 6.
8. A speaker playback system comprising a speaker and a DSP including blocks that apply at least the first-order HRTF, the second-order HRTF and the third-order HRTF, and an adder block that sums the outputs of at least the first-order sound path, the second-order sound path and the third-order sound path for each of the left and right ears, and which creates multiple virtual sound sources using the method of any one of claims 1 to 6.
9. A playback system comprising a pair of stereo speakers, a DSP including a block for applying at least the first-order HRTF, the second-order HRTF, and the third-order HRTF, and an adder block for summing the outputs of at least the first-order sound path, the second-order sound path, and the third-order sound path for the left and right ears, respectively, A system for creating multiple virtual sound sources using the method of any one of claims 1 to 6 to create virtual surround sound sources around a listening room by embedding encoded positional information in the direct sound from the pair of stereo speakers.
10. A gaming engine system that creates a plurality of virtual sound sources using any one of the methods of Claims 1 to 6, comprising headphones or speakers, a DSP including a block for applying at least the primary HRTF, the secondary HRTF and the tertiary HRTF, and an adder block for summing the outputs of at least the primary sound path, the secondary sound path and the tertiary sound path for the left and right ears, respectively, wherein the system creates a plurality of virtual sound sources using any one of Claims 1 to 6.
11. A digital audio workstation (DAW) software system that creates a plurality of virtual sound sources using any one of the methods of Claims 1 to 6, comprising headphones or speakers, a DSP including blocks for applying at least the first-order HRTF, the second-order HRTF and the third-order HRTF, and an adder block for summing the outputs of at least the first-order sound path, the second-order sound path and the third-order sound path for the left and right ears, respectively, wherein the DSP creates a plurality of virtual sound sources using any one of Claims 1 to 6.
Citation Information
Patent Citations
Damiihetsudonyoru supiikasaiseiyorokuonshisutemu
JP1976075401A
JP1982151091U
Video game machine
JP1994198074A