Sound reproduction using multi-order hrtf between left and right ears

By using multi-level HRTF encoding, the problem of unstable sound source localization in headphones is solved, and a realistic surround sound field is reproduced in headphones, which is suitable for virtual reality and low-latency applications.

CN116097664BActive Publication Date: 2026-08-04INNIT AUDIO AB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INNIT AUDIO AB
Filing Date
2021-10-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies cannot effectively locate virtual sound sources outside the listener's head using headphones, and existing methods have poor inter-individual stability, failing to achieve convincing surround sound field reproduction.

Method used

Position encoding is performed using a multi-order head correlation transfer function (HRTF), including at least second-order HRTF encoding between the left and right ears. This interprets the sound source location using time-domain information and creates a realistic surround sound field in headphones.

Benefits of technology

It achieves stable sound source localization in headphones, especially in front of the head, providing a realistic surround sound experience suitable for virtual reality and low-latency applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116097664B_ABST
    Figure CN116097664B_ABST
Patent Text Reader

Abstract

In order to localize sound sources in space, typically the so-called Head Related Transfer Function (HRTF) is applied. Typically, the head related frequency responses (HRFR) for hundreds of individuals are averaged to produce an average HRFR for each position. The average HRFR data is then used for the position coding of audio sources in recording and playback. The invention solves the position coding by decomposing the localization process in a novel way introducing a new time domain focus method. According to the invention, this method is called multi-order HRTF. This method allows averaging across individuals and where its time domain coding provides a more stable sound source localization that is clearly positioned outside the listener's head through headphones. It is also possible to create a virtual surround sound source around the listening room using only two stereo speakers by embedding the encoded position information into the direct sound from the stereo speaker pair.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Introduction In the audio industry, a long-standing goal has been to increase the listener's engagement and immersion in the recorded and subsequently reproduced sound. This exploration was already very active in 1931 when Alan Blumlein invented stereo. Over the years, sound quality, and the resulting immersion, has gradually improved. Although various forms of surround sound existed earlier, Dolby introduced Dolby Stereo in the 1970s, despite its name being the first commercially successful surround sound format. Surround sound offered a higher level of immersion than previously achievable. In recent years, object-based audio formats such as Dolby Atmos and Sony 360 have emerged, further enhancing the level of immersion.

[0002] One of the major challenges associated with all surround sound formats is the reproduction of the surround sound field. While Dolby Atmos commercial cinemas with hundreds of speakers around the screening room sound incredibly impressive, replicating such a setup in a private home is impractical. The industry also struggles to create a convincing reproduction of the surround sound field using headphones. Despite significant research efforts, current technology cannot produce a sound field that is perceived as being distinctly outside the head of a listener wearing headphones. Sound is typically perceived as primarily inside the head and not surrounding the listener as expected. Furthermore, the small amount of sound outside the listener's head is primarily located immediately to the left and right of the listener's ears or slightly behind them. It is impossible to provide a distinctly stable anterior hemisphere position that is highly desirable.

[0003] To locate sound sources in space, a so-called Head-Related Transfer Function (HRTF) is typically applied. HRTF encoding is used for surround sound produced in movies, video games, and many stereo recordings. Location-based HRTF encoding exists in both surround sound and stereo recordings and is applicable to both speaker playback and headphone playback. Several playback algorithms (such as Dolby Atmos for headphones) also employ HRTF encoding to locate sound.

[0004] The research community has released several HRTF databases online, containing measurement results from hundreds of test subjects, and these databases are available for download. These databases typically contain frequency responses associated with multiple locations around each test subject, known as head-related frequency responses (HRFRs). Some databases also include associated time-domain responses, referred to as head-related impulse responses (HRIRs).

[0005] Typically, the HRFR responses of hundreds of individuals are averaged to produce an average HRFR for each location. The average HRFR data is then used for location encoding of the audio source during recording and playback.

[0006] As previously discussed, this type of averaging HRFR coding does not produce convincing results with headphones and requires multiple speakers distributed around the room. Even when measurements are averaged across many test subjects, the perceived location varies significantly between individuals.

[0007] However, successful results can be obtained by using individually measured HRIR for each listener. Convolving playback material with individual HRIR using a standard FIR filter can create a fully realistic surround sound immersion with headphones, but this only works for individuals using their personal HRIR during playback convolution. Generating individual HRIR data for each person listening to the recording is clearly not feasible. Numerous attempts have been made to customize commonly used average HRIR data based on information provided by individuals regarding their personal physical characteristics, but without any breakthroughs.

[0008] For HRIR, the latency in the FIR filter also becomes a problem. To obtain good results, HRIR must be quite long, and the introduced latency will cause significant problems in virtual reality, gaming, and other similar applications, where significant latency is unacceptable.

[0009] In the time domain, successful direct averaging methods such as HRFR averaging are also not possible. Figure 1 This illustrates the difficulty of time-domain HRIR averaging. Figure 1 Traces 1, 2, and 3 in the diagram illustrate HRIR data from three different test subjects. Due to the different physical sizes and associated acoustic travel times, the second bulge in the HRIR data occurs at different time points relative to the larger first arrival on the left side of the trace. Traces 4 show the average of 1, 2, and 3. Clearly, this is not a good average from three physically different test subjects. In this example, trace 2 would be the best average between the individual sizes, but trace 4 would not look exactly like trace 2. The three individual bulges on traces 1 through 3 have become temporally blurred. Instead of a clear wavefront arrival at the average time point (trace 2), the wavefront has been blurred and suppressed by time, which is not the desired result.

[0010] This invention addresses location encoding by decomposing the localization process in a novel way that introduces a new temporal focus method. This method is called multi-order HRTF. It allows averaging across individuals, and its temporal encoding provides more stable audio source localization, which, if desired, is clearly located outside and in front of the listener's head via headphones. Virtual surround sound sources can also be created around the listening room using only two stereo speakers by embedding the encoded location information into the direct sound from a stereo speaker pair. Summary of the Invention

[0011] This invention relates to a method for sound reproduction, the method comprising positional encoding using a multi-order head correlation transfer function (HRTF), wherein the method involves using at least a first-order HRTF from the left ear, and then using a second-order HRTF from the left ear to the right ear, while simultaneously using a first-order HRTF from the right ear, and then using a second-order HRTF from the right ear to the left ear, thereby reproducing the sound. Regarding the foregoing, it should be noted that "multi-order" can mean second-order, third-order, or up to any order. It can also be mentioned that, according to one embodiment, the method involves at least a third-order HRTF from left ear to right ear in the same manner as from right ear to left ear, preferably involving at least a fourth-order HRTF from left ear to right ear in the same manner as from right ear to left ear.

[0012] Furthermore, in the following text and regarding the accompanying drawings (especially regarding...) Figure 2 The concept according to the invention is further described below.

[0013] Furthermore, regarding this invention, it should be mentioned that many known methods exist that use several or more HRTFs, such as those disclosed in US2020 / 0037097; however, these differ from the concepts disclosed and provided in this invention. Moreover, this invention provides a method comprising using at least a first-order HRTF for the left ear, followed by a second-order HRTF from the left ear to the right ear, while simultaneously using a first-order HRTF for the right ear, and then using a second-order HRTF from the right ear to the left ear, thereby reproducing sound. This should not be confused with the use of several or more HRTFs utilized in many known methods. Attached Figure Description

[0014] Figure 1 This illustrates the difficulty of time-domain HRIR averaging.

[0015] Figure 2 It shows the sound path from the sound source to the listener's head and around that head.

[0016] Figure 3 It shows the relationship with Figure 2 The frequency response associated with sound position 2 and sound path 6 in the sound.

[0017] Figure 4 A block diagram containing a typical multi-stage HRTF DSP implementation.

[0018] Detailed description of multi-level HRTF As is well known from psychoacoustic research, human hearing is extremely sensitive to the temporal characteristics of sound. Differences in sound waves between wood and metal can be heard within the first few milliseconds after a material is struck. The starting waveforms of violin and trumpet notes are very different, and the differences are easily heard. However, if sustained notes from each instrument are heard without any initial attack, it becomes difficult to distinguish between them.

[0019] In the same way, sound source location is interpreted not only by HRFR but also by time-domain information. Due to the difficulties discussed, previous localization solutions focused on average HRFR data while ignoring time-domain information. The results were less than convincing. Individual HRIR data captures time-domain information, but only for one individual at a time, and manages to provide a good surround sound field impression for the individual in question.

[0020] Figure 2 The diagram shows the sound path from the sound source to the listener's head and around that head. Number 1 is the listener, number 2 is the sound source, and numbers 3 through 8 are the visualized sound wave paths to and around the head. Figure 2 Only one sound source location is shown, but any location in three-dimensional space has a similar set of imaginable sound paths associated with it. Figure 2 It shows that the general principles and paths for other sound source locations should be easily extrapolated.

[0021] Each sound path 3 through 8 has its associated time delay, frequency response, and attenuation. Path 3 has a time delay, but in this particular case, it has the travel time of the sound from source 2 to the right ear. Since this is the first time the sound reaches the listener, the delay is zero because there is no need for a delay parallel to the sound's travel time to reach the listener. Since the sound travels directly to the ear without any obstacles that could cause attenuation, the attenuation in this specific first-order path is also zero. The frequency response will generally be the well-known average HRFR of the source location relative to the right ear. However, the sound wave will not stop once it has reached the right ear. It will continue along path 6 around the head to the left ear. This path has an interaural time delay due to the sound's travel time, a frequency response due to the masking of higher frequencies by the head, etc., and attenuation caused by traveling around the head to the other ear. This second-wave path is a second-order HRTF. When the sound wave has reached the left ear, it will again continue along path 8 back to the right ear, and again, this path has its associated time delay, frequency response, and attenuation. This is a third-order HRTF. For clarity, Figure 2 Higher-order HRTFs are not shown, but the principle should now be obvious and any higher-order HRTFs can be easily extrapolated simply by continuing the path around the head.

[0022] The time delay associated with the path between ears is directly related to the physical distance between them, and this time delay is approximately 200 µs to 1 ms, typically around 600 µs. As sound waves travel across the head from one ear to the other, the change in frequency response caused by that head movement is typically higher in the spectrum, starting at 400 Hz to 2.5 kHz and sloping downwards until reaching the human hearing limit of 20 kHz and above. Due to the physical characteristics of the human head and shoulders, several troughs and peaks may exist associated with a specific path. Attenuation typically varies from 0 to 6 dB in a first-order path, from 3 dB to 12 dB in a second-order path, from 6 dB to 24 dB in a third-order path, and from 9 dB to 48 dB in a fourth-order path. For those skilled in the art using standard methods, the methods and techniques involved in obtaining the exact time delay and attenuation associated with each path should be straightforward and are therefore not discussed further.

[0023] The frequency response involved can be determined from readily available HRTF data. Figure 3 It shows the relationship with Figure 2 The frequency response associated with sound position 2 and sound path 6, i.e. amplitude (dB) versus frequency (Hz).

[0024] Acoustic measurements have shown that, as described above, the sound waves propagate several times around the object, and when second-order, third-order, and fourth-order HRTFs are added, the sound is perfectly audible and clear, perceived as more natural, and the localization of the sound source is greatly improved. Localization and naturalness improve with increasing orders up to the fourth, after which the improvement becomes less noticeable. Of course, any order of HRTF, from second to as many as one can imagine (hundreds or even thousands), can be used, but as mentioned above, orders above the fourth offer only minor benefits.

[0025] Similar to the path described above that begins with path 3, the sound path that begins with path 4 from the initial sound source to the left ear also has the time delay, frequency response, and attenuation associated with each of them. However, the delay along path 4 is not zero like that along path 3; there is a delay due to the interaural time difference. The frequency change that occurs will typically again be the well-known average HRFR of the sound source location relative to the left ear. The attenuation along path 4 is typically 4.5 dB, where the sound source is located as shown in the example. The subsequent second-order path 5 and the subsequent third-order path 7 also have associated time delays, frequency responses, and attenuations.

[0026] The sound path across the front of the head from one ear to the other is slightly longer than the path across the back of the head. This sound path also produces slightly different attenuation and frequency shifts compared to the sound path across the back of the head. Considering this, it becomes apparent that the head and ears are excellent localization devices, where different sound source locations will produce unique sets of multi-order HRTF sound paths. Therefore, multi-order HRTFs make it possible to achieve stable localization of sound sources both in front of and behind the head.

[0027] With multi-order HRTFs separating the frequency response changes, the attenuation and time delay used to average each path across the test subject become straightforward. Familiar methods can be used to easily average the frequency response across many individuals for each path, and the attenuation and delay become the average of the attenuation and travel distance for each test subject across only each path. Averaging the characteristics of many individuals is crucial for achieving stable and similar results for all listeners.

[0028] Frequency changes associated with each path can be easily implemented using standard IIR filters that eliminate the delay associated with FIR filters. Therefore, multi-order HRTFs operate without introducing any delay, making this approach ideal for virtual reality, gaming, and any other application requiring zero or extremely low latency. Figure 4 A block diagram containing a typical multi-order HRTFDSP implementation is provided. A fourth-order implementation for a single sound source location is shown. It is certainly possible and apparent that multi-order HRTF can be implemented in many other ways, and Figure 4 Only one example of a topology from many possible topologies is shown. Blocks 11, 21, 31, 41, 51, 61, 71, and 81 are delay blocks in a fourth-order implementation that apply delays associated with each of the four paths for each ear. Blocks 12, 22, 32, 42, 52, 62, 72, and 82 apply frequency changes associated with each path. Blocks 13, 23, 33, 43, 53, 63, 73, and 83 are gain blocks that apply attenuation present in each path. Finally, 100 is an adder block that simply sums all outputs from the four paths to the left ear, and 200 is an adder for the right ear. The outputs from 100 and 200 are sent to the corresponding left and right channels.

[0029] Multi-level HRTF applications can handle both stereo and multi-channel input signals. Multiple virtual sound sources can be created using multi-level HRTF. If the input signal is in a standard five-channel surround sound format, multi-level HRTF can be used to create five virtual speakers positioned in the typical locations of a five-channel surround sound setup (i.e., front left and front right, center, and left and right surround). The discrete input channels are then played back by their respective virtual speakers. Similarly, more virtual speakers can be created for the latest surround sound formats involving more surround speakers and additional ceiling speakers. For stereo input signals, a standard sound extraction and guide process can be used to extract the individual feeds to the virtual speakers. In this case, the stereo extraction and guide process will be the same as for standard surround sound products.

[0030] Virtual sound sources created using multi-order HRTF work for both headphones and speakers. With headphones, a surround sound field can be created that approximates the experience of using individually measured HRIR. For speakers, virtual speakers can be encoded as direct sound from a pair of stereo speakers that create virtual center, surround, and height speakers. Using multi-order HRTF virtual speakers, a surround sound field can be created that is perceived as similar to a setup with multiple speakers.

[0031] Playback using multi-level HRTF virtual sound sources is certainly not limited to today's stereo or surround formats and their source locations. The examples above only illustrate possible multi-level HRTF applications, and of course, any number of virtual speakers can be created in any location as needed.

[0032] Multi-order HRTF can be applied at any stage from sound recording / generation to playback, not just the playback stage. Multi-order HRTF can be used in design and / or production to apply location to sound, which can later be played back on headphones, standard stereo, or multi-channel playback systems. As an example, multi-order HRTF can be used within a game engine to locate sounds within the generated game sound field. Another example is using multi-order HRTF (integrated or as a plugin) within DAW software to locate sounds within a sound field during sound production. In other words, multi-order HRTF algorithms and sound processing can be applied at any stage that provides the same end result.

[0033] Detailed Implementation Plan The following section provides some specific embodiments of the present invention.

[0034] According to one embodiment of the invention, the method includes at least a three-order HRTF from left ear to right ear in the same manner as from right ear to left ear, preferably including at least a four-order HRTF from left ear to right ear in the same manner as from right ear to left ear.

[0035] Furthermore, according to another embodiment, the method includes creating one or more virtual sound sources by embedding encoded location information into the sound.

[0036] According to another implementation, each head-related transfer function (HRTF) from second-order and above includes parametric time delay, frequency response, and attenuation.

[0037] Furthermore, according to another embodiment, the method takes into account the differences in different sound paths, such as the difference between a sound path from one ear to the other in front of the head and a sound path at the back of the head. In this respect, it should be noted that a sound path from one ear to the other can be any path around the head. Therefore, the method according to the invention can involve several sound paths.

[0038] Furthermore, according to yet another embodiment, the method includes averaging. As disclosed above, according to the invention, averaging can be performed across individuals. Utilizing time-domain coding provides more stable sound source localization, whereby the sound source is clearly located outside and in front of the listener's head if desired. Based on this, according to one embodiment of the invention, the method includes averaging with a focus on the time domain. Furthermore, according to one embodiment of the invention, the method includes averaging parameters such as time delay, frequency response, and attenuation independently of each other. This is yet another difference when compared to averaging performed in known methods currently in use.

[0039] This invention also relates to different types of system, hardware, and software implementations.

[0040] According to one embodiment, the present invention relates to a headphone playback system arranged for use with the method according to the invention.

[0041] Furthermore, the present invention also relates to a loudspeaker playback system arranged for use with the method according to the present invention.

[0042] Furthermore, the present invention relates to a playback system comprising a pair of stereo speakers, the system being arranged to use the method according to the invention to create a virtual surround sound source around a listening room by embedding coded location information into the direct sound from the pair of stereo speakers.

[0043] According to the present invention, other applications are also possible, as can be clearly seen from the above description.

[0044] According to one such embodiment, the present invention relates to a game engine system configured to use the method according to the invention. According to another embodiment, the present invention provides a digital audio workstation (DAW) software system configured to use the method according to the invention.

Claims

1. A method for sound reproduction, the method comprising positional encoding using a multi-order head correlation transfer function (HRTF), wherein the method involves using at least a first-order HRTF from the left ear, and then using a second-order HRTF from the left ear to the right ear, while simultaneously using a first-order HRTF from the right ear, and then using a second-order HRTF from the right ear to the left ear, thereby reproducing the sound, wherein the method comprises using at least a third-order HRTF from the left ear to the right ear in the same manner as from the right ear to the left ear.

2. The method of claim 1, wherein the method includes at least a fourth-order HRTF from the left ear to the right ear in the same manner as from the right ear to the left ear.

3. The method according to claim 1 or 2, wherein the method includes creating one or more virtual sound sources by embedding encoded location information into the sound.

4. The method according to any one of claims 1 to 3, wherein each head-related transfer function (HRTF) from the second order and above includes parametric time delay, frequency response, and attenuation.

5. The method according to any one of claims 1 to 4, wherein the method takes into account the differences between different sound paths, including the difference between the sound path from one ear to the other in front of the head and the sound path behind the head.

6. The method according to any one of claims 1 to 5, wherein the method comprises averaging the parameter time delay, frequency response, and attenuation independently of each other.

7. The method according to any one of claims 1 to 6, wherein the method includes averaging with a focus on the time domain.

8. A headphone playback system, the headphone playback system being arranged for use with the method according to any one of claims 1 to 7.

9. A loudspeaker playback system, the loudspeaker playback system being arranged for use with the method according to any one of claims 1 to 7.

10. A playback system comprising a pair of stereo speakers, the system being arranged to use the method according to any one of claims 1 to 7 to create a virtual surround sound source around a listening room by embedding coded location information into the direct sound from the pair of stereo speakers.

11. A game engine system configured to use the method according to any one of claims 1 to 7.

12. A digital audio workstation (DAW) software system, said DAW software system being configured to use the method according to any one of claims 1 to 7.